OnCallReady

SRE · 7 min read

How to run an incident as incident commander: roles, severity, cadence and handoff

The incident commander does not debug. A real-shaped SEV2, from four people changing production at once to a clean handoff, and what the IC says at each step.

You join the channel at 18:55. Checkout has been failing since 18:38. Someone acknowledged the page at 18:41, and since then the room looks like this:

terminal
$ incident log
18:55  bot       PagerDuty: [FIRING] CheckoutErrorBudgetBurn severity=page (acknowledged by @mo 18:41)
18:55  bot       #checkout-dev is already busy. You are the on-call SRE: you join at 18:55.
18:55  @mo       restarting the checkout pods again, that helped for a bit last time
18:55  @dana     I am raising the DB pool from 20 to 50 on checkout, maybe it is connections
18:55  @lee      scaling checkout to 12, CPU looks a bit high on two pods
18:55  @priya    who just doubled the connections on orders-db?? I see 380/400 used
18:55  @mo       restart done. errors went down. no wait they are back

Four capable people, four changes to production, nobody knows what anyone else did, nobody has asked what changed, and support has 31 tickets and no answer. This is "hero mode", and it is the most common way a 20-minute incident becomes a two-hour one. (The incident command is the OnCallReady incident room, a simulator. The room, the people and the clock are made up. The moves are the real ones from the Google SRE book, PagerDuty's incident response guide and Atlassian's handbook.)

What an incident commander actually does

The incident commander (IC) owns the incident, not the fix. Google's SRE book sums the job up as the three Cs: coordinate, communicate, control. The IC sets the severity, hands out roles, decides (roll back or not, page more people or not, when it is over) and keeps everyone informed. The IC does not debug. The moment the IC opens a terminal, nobody is watching the whole room.

The roles, by their common names:

text
IC           owns the incident: severity, roles, decisions, cadence, when it ends
ops lead     the only hands on production; investigates and mitigates (PagerDuty: SMEs, Atlassian: tech lead)
comms lead   writes the updates to every audience (PagerDuty: customer + internal liaison)
scribe       keeps the timeline: what was seen, decided and done, with times

Small incident: one person holds every role. When a second person arrives, split IC from ops first, then comms, then scribe. Any role not handed out belongs to the IC.

1. Declare, at a severity, with a one-line summary

Severity is a function of impact, not of how scary the graph looks or how senior the person asking is. A matrix makes it a lookup instead of a debate:

terminal
$ incident matrix
Impact matrix (the lab's severity policy; every company writes its own)

                            most users  a subset    a few
                            (>50%)      (5-50%,     (<5%, one
                                        a region)   customer)
  core journey down         SEV1        SEV2        SEV3
  core journey degraded     SEV2        SEV3        SEV3
  non-core or internal      SEV3        SEV3        SEV4
  no user impact yet        SEV4        SEV4        SEV4

  Any data loss, security breach or money moved wrongly: SEV1, whatever the numbers.
  About to hit users within the hour? Judge by the impact it is about to have.

Update cadence: SEV1 every 30 min, SEV2 every 30 min, SEV3 every 60 min, SEV4 none (a ticket).

A third of customers cannot pay: core journey down for a subset, SEV2. Declare early: a severity can go down later, and an undeclared incident has no commander at all.

terminal
$ incident declare --sev 2 "checkout failing for about a third of customers since 18:38"
Declared INC-4502 SEV2 at 18:55 UTC.
18:55  bot       INC-4502 declared SEV2 by @learner - you are incident commander until you hand it on. Channel #inc-4502-checkout. Policy: updates every 30 min (first due 19:25 UTC).

2. Stop the uncoordinated changes

The first command, to everyone, in plain words:

terminal
$ incident say "Everyone: stop all changes to production now. Nobody touches prod unless I assign it. I am IC."
18:58  @mo       ok, no more restarts from me.
18:58  @lee      Stopping. Scale stays at 12, I am hands off.

Then one ops lead, and a scribe so the timeline survives:

terminal
$ incident role ops @dana
OPS: @dana
19:01  bot       @dana is now ops lead (the only hands on production)
$ incident role scribe @lee
SCRIBE: @lee
19:02  bot       @lee is now scribe

"One pair of hands on production" is the rule that stops two fixes from colliding. In this room, Dana's pool change had already pushed orders-db to 380 of its 400 connections.

3. Ask what changed, then decide with a reason

Most incidents follow a change. Nobody had asked:

terminal
$ incident say "@mo what changed today?"
19:04  @mo       2.8.4 went out at 18:36 - it adds a call to the new currency-rates service, behind the feature flag fx-v2. That is the only change today. Rolling back to 2.8.3 is safe (no schema change).

A recent change plus a safe rollback means you roll back first and understand it later. Mitigate before you diagnose. Record the decision with its reason, and give the task to one person, by name:

terminal
$ incident decide "roll back checkout to 2.8.3" --why "errors started two minutes after 2.8.4 and Mo says the rollback is safe"
Decision recorded at 19:04 UTC.
$ incident assign @dana "roll back checkout to 2.8.3, nobody else touches prod"
Task #1 -> @dana
19:06  @dana     On it, will report back.

4. Talk to every audience, on a clock

Each update has three parts: impact in the reader's words, what we are doing, and when the next update comes. Never a fix time. The details per audience are in how to write an incident status update. Support first, because they have 31 customers waiting:

terminal
$ incident update support "Customers who cannot pay at checkout: it is a problem on our side and we are fixing it now. They can try again in about 15 minutes; PayPal may work meanwhile. Next update at 19:30 UTC."
Posted the support update at 19:06 UTC; next update due 19:30 UTC.
$ incident update status --component partial "Some customers cannot complete payment at checkout. We have found the cause and are fixing it. Next update by 19:30 UTC."
Posted the status update at 19:09 UTC; next update due 19:30 UTC.

Then keep the cadence (every 30 minutes for a SEV2 here), even when nothing changed. "No change, still rolling back, next update 19:55" is a complete update. Silence makes people escalate around you.

5. Resolve on evidence, then hand it to the postmortem

The rollback finishes at 19:11 and errors drop. "Errors are dropping" is not "resolved". Wait for the metric to be healthy for a set time (15 minutes in this policy), post the final status update, then close:

terminal
$ incident resolve "2.8.4 rolled back; checkout healthy since 19:15; uncoordinated changes stopped at 18:58"
Resolved INC-4502.
19:44  bot       INC-4502 resolved by @learner at 19:44 UTC - 49 min after it was declared. Metric healthy for 31 min.
$ incident postmortem --owner @lee --date 2026-09-18
Postmortem scheduled: @lee, 2026-09-18.

The scribe's timeline is the raw material for the postmortem. Who decided what, and when, is gone by the next morning if nobody wrote it down.

Handing off command

Incidents outlast shifts. A handoff is when information gets lost, so it has a fixed form: five parts, an explicit acceptance, and an announcement. The room's reviewer checks a draft the way a good incoming IC would:

terminal
$ incident check handoff "Emails are slow, provider issue. Dana is on it. Good luck!"
handoff draft (simulator - the five things the next IC needs):
  impact right now:                      yes
  what we know (hypothesis, ruled out):  MISSING
  in flight, and who owns it (@name):    yes
  next update due (a time):              MISSING
  open decisions or risks:               MISSING
Add: what we know so far (the current hypothesis, what is ruled out); when the next update is due (a time); the open decisions or risks.

A complete one, for an email outage at the end of a European shift:

text
Impact now: order confirmation emails are about 2 hours late for most customers; orders and
payments are fine. What we know: the email provider throttled us after the 14:00 marketing send;
they raised us to 200/min. In flight: @dana is draining the 38,000-email queue (empty around
19:50), @nina is asking the provider for 500/min (ticket 88213). Next update due 16:55 UTC.
Open decision: the 19:00 marketing send - if it goes, we are throttled again; I recommend
postponing it.

The new IC confirms in their own words (I have IC. Confirming: SEV2, next update due 16:55 UTC.), and the channel hears it: "IC is now @casey". Google's book is explicit that the outgoing commander does not leave until the handoff has been acknowledged. And the next update still goes out on time. A handoff is no excuse for a gap.

When someone more senior joins and starts giving orders, PagerDuty's answer is a question asked in the open: "Do you wish to take command?" If yes, do this handoff. If no, they get the executive update like everyone else.

The loop, in one screen

Between those big moments the IC runs the same short loop, again and again: size up (impact now, who is here), stabilize (anyone changing things uncoordinated? is the best mitigation owned and time-boxed?), update (has every audience heard from us on time?), verify (did the last action do what we expected? look at the impact, not the code). It works the same for a sealed Vault taking down every service as for a latency incident with no errors at all.

Practise it

The chapter Incident Command & Communication runs all of this in a simulated incident room with a pager, a status page and a clock. Responders answer what you post, and support chases you when an update is late. Start with the mission Run a SEV2 from page to resolve, then Incident: everyone is debugging, nobody is talking (the room above) and Hand off command at the end of your shift. Each lab checks the declare time, the roles, the decisions and every update you send.

OnCallReady is free, with no ads and no tracking. RSS · All posts