You join the channel at 18:55. Checkout has been failing since 18:38. Someone acknowledged the page at 18:41, and since then the room looks like this:
$ incident log
18:55 bot PagerDuty: [FIRING] CheckoutErrorBudgetBurn severity=page (acknowledged by @mo 18:41)
18:55 bot #checkout-dev is already busy. You are the on-call SRE: you join at 18:55.
18:55 @mo restarting the checkout pods again, that helped for a bit last time
18:55 @dana I am raising the DB pool from 20 to 50 on checkout, maybe it is connections
18:55 @lee scaling checkout to 12, CPU looks a bit high on two pods
18:55 @priya who just doubled the connections on orders-db?? I see 380/400 used
18:55 @mo restart done. errors went down. no wait they are backFour capable people, four changes to production, nobody knows what anyone else did, nobody has asked what changed, and support has 31 tickets and no answer. This is "hero mode", and it is the most common way a 20-minute incident becomes a two-hour one. (The incident command is the OnCallReady incident room, a simulator. The room, the people and the clock are made up. The moves are the real ones from the Google SRE book, PagerDuty's incident response guide and Atlassian's handbook.)
What an incident commander actually does
The incident commander (IC) owns the incident, not the fix. Google's SRE book sums the job up as the three Cs: coordinate, communicate, control. The IC sets the severity, hands out roles, decides (roll back or not, page more people or not, when it is over) and keeps everyone informed. The IC does not debug. The moment the IC opens a terminal, nobody is watching the whole room.
The roles, by their common names:
IC owns the incident: severity, roles, decisions, cadence, when it ends
ops lead the only hands on production; investigates and mitigates (PagerDuty: SMEs, Atlassian: tech lead)
comms lead writes the updates to every audience (PagerDuty: customer + internal liaison)
scribe keeps the timeline: what was seen, decided and done, with timesSmall incident: one person holds every role. When a second person arrives, split IC from ops first, then comms, then scribe. Any role not handed out belongs to the IC.
1. Declare, at a severity, with a one-line summary
Severity is a function of impact, not of how scary the graph looks or how senior the person asking is. A matrix makes it a lookup instead of a debate:
$ incident matrix
Impact matrix (the lab's severity policy; every company writes its own)
most users a subset a few
(>50%) (5-50%, (<5%, one
a region) customer)
core journey down SEV1 SEV2 SEV3
core journey degraded SEV2 SEV3 SEV3
non-core or internal SEV3 SEV3 SEV4
no user impact yet SEV4 SEV4 SEV4
Any data loss, security breach or money moved wrongly: SEV1, whatever the numbers.
About to hit users within the hour? Judge by the impact it is about to have.
Update cadence: SEV1 every 30 min, SEV2 every 30 min, SEV3 every 60 min, SEV4 none (a ticket).A third of customers cannot pay: core journey down for a subset, SEV2. Declare early: a severity can go down later, and an undeclared incident has no commander at all.
$ incident declare --sev 2 "checkout failing for about a third of customers since 18:38"
Declared INC-4502 SEV2 at 18:55 UTC.
18:55 bot INC-4502 declared SEV2 by @learner - you are incident commander until you hand it on. Channel #inc-4502-checkout. Policy: updates every 30 min (first due 19:25 UTC).2. Stop the uncoordinated changes
The first command, to everyone, in plain words:
$ incident say "Everyone: stop all changes to production now. Nobody touches prod unless I assign it. I am IC."
18:58 @mo ok, no more restarts from me.
18:58 @lee Stopping. Scale stays at 12, I am hands off.Then one ops lead, and a scribe so the timeline survives:
$ incident role ops @dana
OPS: @dana
19:01 bot @dana is now ops lead (the only hands on production)
$ incident role scribe @lee
SCRIBE: @lee
19:02 bot @lee is now scribe"One pair of hands on production" is the rule that stops two fixes from colliding. In this room, Dana's pool change had already pushed orders-db to 380 of its 400 connections.
3. Ask what changed, then decide with a reason
Most incidents follow a change. Nobody had asked:
$ incident say "@mo what changed today?"
19:04 @mo 2.8.4 went out at 18:36 - it adds a call to the new currency-rates service, behind the feature flag fx-v2. That is the only change today. Rolling back to 2.8.3 is safe (no schema change).A recent change plus a safe rollback means you roll back first and understand it later. Mitigate before you diagnose. Record the decision with its reason, and give the task to one person, by name:
$ incident decide "roll back checkout to 2.8.3" --why "errors started two minutes after 2.8.4 and Mo says the rollback is safe"
Decision recorded at 19:04 UTC.
$ incident assign @dana "roll back checkout to 2.8.3, nobody else touches prod"
Task #1 -> @dana
19:06 @dana On it, will report back.4. Talk to every audience, on a clock
Each update has three parts: impact in the reader's words, what we are doing, and when the next update comes. Never a fix time. The details per audience are in how to write an incident status update. Support first, because they have 31 customers waiting:
$ incident update support "Customers who cannot pay at checkout: it is a problem on our side and we are fixing it now. They can try again in about 15 minutes; PayPal may work meanwhile. Next update at 19:30 UTC."
Posted the support update at 19:06 UTC; next update due 19:30 UTC.
$ incident update status --component partial "Some customers cannot complete payment at checkout. We have found the cause and are fixing it. Next update by 19:30 UTC."
Posted the status update at 19:09 UTC; next update due 19:30 UTC.Then keep the cadence (every 30 minutes for a SEV2 here), even when nothing changed. "No change, still rolling back, next update 19:55" is a complete update. Silence makes people escalate around you.
5. Resolve on evidence, then hand it to the postmortem
The rollback finishes at 19:11 and errors drop. "Errors are dropping" is not "resolved". Wait for the metric to be healthy for a set time (15 minutes in this policy), post the final status update, then close:
$ incident resolve "2.8.4 rolled back; checkout healthy since 19:15; uncoordinated changes stopped at 18:58"
Resolved INC-4502.
19:44 bot INC-4502 resolved by @learner at 19:44 UTC - 49 min after it was declared. Metric healthy for 31 min.
$ incident postmortem --owner @lee --date 2026-09-18
Postmortem scheduled: @lee, 2026-09-18.The scribe's timeline is the raw material for the postmortem. Who decided what, and when, is gone by the next morning if nobody wrote it down.
Handing off command
Incidents outlast shifts. A handoff is when information gets lost, so it has a fixed form: five parts, an explicit acceptance, and an announcement. The room's reviewer checks a draft the way a good incoming IC would:
$ incident check handoff "Emails are slow, provider issue. Dana is on it. Good luck!"
handoff draft (simulator - the five things the next IC needs):
impact right now: yes
what we know (hypothesis, ruled out): MISSING
in flight, and who owns it (@name): yes
next update due (a time): MISSING
open decisions or risks: MISSING
Add: what we know so far (the current hypothesis, what is ruled out); when the next update is due (a time); the open decisions or risks.A complete one, for an email outage at the end of a European shift:
Impact now: order confirmation emails are about 2 hours late for most customers; orders and
payments are fine. What we know: the email provider throttled us after the 14:00 marketing send;
they raised us to 200/min. In flight: @dana is draining the 38,000-email queue (empty around
19:50), @nina is asking the provider for 500/min (ticket 88213). Next update due 16:55 UTC.
Open decision: the 19:00 marketing send - if it goes, we are throttled again; I recommend
postponing it.The new IC confirms in their own words (I have IC. Confirming: SEV2, next update due 16:55 UTC.), and the channel hears it: "IC is now @casey". Google's book is explicit that the outgoing commander does not leave until the handoff has been acknowledged. And the next update still goes out on time. A handoff is no excuse for a gap.
When someone more senior joins and starts giving orders, PagerDuty's answer is a question asked in the open: "Do you wish to take command?" If yes, do this handoff. If no, they get the executive update like everyone else.
The loop, in one screen
Between those big moments the IC runs the same short loop, again and again: size up (impact now, who is here), stabilize (anyone changing things uncoordinated? is the best mitigation owned and time-boxed?), update (has every audience heard from us on time?), verify (did the last action do what we expected? look at the impact, not the code). It works the same for a sealed Vault taking down every service as for a latency incident with no errors at all.
Practise it
The chapter Incident Command & Communication runs all of this in a simulated incident room with a pager, a status page and a clock. Responders answer what you post, and support chases you when an update is late. Start with the mission Run a SEV2 from page to resolve, then Incident: everyone is debugging, nobody is talking (the room above) and Hand off command at the end of your shift. Each lab checks the declare time, the roles, the decisions and every update you send.