OPERATIONS GUIDE · OPS-200

Operate the control plane through explicit service loops.

A practical operating model for routine review, anomaly triage, controlled change, incident coordination, maintenance, and evidence-based handoff.

Owner
Global Reliability
Version
1.3
Reviewed
05 Aug 2026
Classification
Public

1. Operating model

NBS separates observation, decision, execution, and verification. The operator who observes a condition may propose a change, but authority is derived from role and the associated change record. Every meaningful transition should have a named owner and a measurable postcondition.

LOOP-01

Observe

Review current state, time range, data freshness, and affected scope.

LOOP-02

Decide

Choose the least disruptive action and state its expected result.

LOOP-03

Execute

Perform one authorized transition and capture its request identifier.

LOOP-04

Verify

Compare observed state with the declared success and rollback criteria.

2. Routine service review

Start each review with active incidents, unresolved high alerts, scheduled changes, regional health, transfer backlog, and capacity exceptions. Confirm the time basis of every indicator. A dashboard value without a current timestamp is supporting context, not decisive evidence.

CadenceReviewOutput
ContinuousHigh-severity alerts and active incidentsOwner and next update time
Per shiftRegional health, blocked transfers, pending changesHandoff record
WeeklyCapacity trend, repeated alerts, access reviewsImprovement actions
QuarterlyControl evidence and policy exceptionsControl review decision

3. Triage an abnormal condition

  1. Define the earliest observed time and affected region, pipeline, transfer, or identity boundary.
  2. Check whether the condition is already represented by an incident or approved maintenance.
  3. Separate symptom from likely cause; do not change several control variables at once.
  4. Assign severity from confirmed impact, not from graph appearance alone.
  5. Escalate when authority, evidence, or containment scope exceeds the current role.
Stop condition

If a security boundary may be involved, preserve request IDs and audit evidence, stop speculative changes, and enter the security response workflow.

4. Execute a controlled change

A change record should identify scope, owner, risk, maintenance window, validation, rollback, and communication. Immediately before execution, re-check current state and conflicting changes. After execution, verify the smallest set of indicators that proves the intended state.

Precondition → approved scope and stable baseline · Execution → one bounded action · Verification → state, health, and audit record · Closure → evidence attached and owner notified

5. Incident coordination

The incident commander owns severity, timeline, decision cadence, and communications. Technical owners investigate within assigned workstreams. Material changes are linked to the incident and executed through the normal authorization boundary unless emergency authority is explicitly invoked and later reviewed.

6. Shift handoff

A handoff includes active incidents, unresolved high alerts, changes in progress, maintenance windows, blocked transfers, temporary mitigations, evidence locations, named owners, and the next decision time. State what changed, what remains uncertain, and what action is expected next.