OPERATIONS GUIDE · OPS-200
Operate the control plane through explicit service loops.
A practical operating model for routine review, anomaly triage, controlled change, incident coordination, maintenance, and evidence-based handoff.
1. Operating model
NBS separates observation, decision, execution, and verification. The operator who observes a condition may propose a change, but authority is derived from role and the associated change record. Every meaningful transition should have a named owner and a measurable postcondition.
Observe
Review current state, time range, data freshness, and affected scope.
Decide
Choose the least disruptive action and state its expected result.
Execute
Perform one authorized transition and capture its request identifier.
Verify
Compare observed state with the declared success and rollback criteria.
2. Routine service review
Start each review with active incidents, unresolved high alerts, scheduled changes, regional health, transfer backlog, and capacity exceptions. Confirm the time basis of every indicator. A dashboard value without a current timestamp is supporting context, not decisive evidence.
| Cadence | Review | Output |
|---|---|---|
| Continuous | High-severity alerts and active incidents | Owner and next update time |
| Per shift | Regional health, blocked transfers, pending changes | Handoff record |
| Weekly | Capacity trend, repeated alerts, access reviews | Improvement actions |
| Quarterly | Control evidence and policy exceptions | Control review decision |
3. Triage an abnormal condition
- Define the earliest observed time and affected region, pipeline, transfer, or identity boundary.
- Check whether the condition is already represented by an incident or approved maintenance.
- Separate symptom from likely cause; do not change several control variables at once.
- Assign severity from confirmed impact, not from graph appearance alone.
- Escalate when authority, evidence, or containment scope exceeds the current role.
If a security boundary may be involved, preserve request IDs and audit evidence, stop speculative changes, and enter the security response workflow.
4. Execute a controlled change
A change record should identify scope, owner, risk, maintenance window, validation, rollback, and communication. Immediately before execution, re-check current state and conflicting changes. After execution, verify the smallest set of indicators that proves the intended state.
Precondition → approved scope and stable baseline · Execution → one bounded action · Verification → state, health, and audit record · Closure → evidence attached and owner notified5. Incident coordination
The incident commander owns severity, timeline, decision cadence, and communications. Technical owners investigate within assigned workstreams. Material changes are linked to the incident and executed through the normal authorization boundary unless emergency authority is explicitly invoked and later reviewed.
6. Shift handoff
A handoff includes active incidents, unresolved high alerts, changes in progress, maintenance windows, blocked transfers, temporary mitigations, evidence locations, named owners, and the next decision time. State what changed, what remains uncertain, and what action is expected next.
Operational