Why this chapter exists
You already know the words. Chapter 0 (0.29) taught severity, the incident commander, mitigating first and the update cadence; lesson 29.15 turned them into the first ten minutes of an incident and a postmortem template. Every one of those lessons had a quiet assumption: that the people involved do what the process says.
They do not. At 18:55 on a Tuesday four engineers are changing production at the same time, each convinced they are helping. The support lead is asking "is something going on?" in three channels. A manager wants a one-liner for the executives. Nobody has declared anything, so nobody is in charge, so nobody can say "stop". The bug in that incident takes ten minutes to fix. The incident lasts an hour.
This chapter is about the people half of an incident: taking command of a room that is already in motion, saying the right thing to six audiences on a clock, paging the right person and escalating when they do not answer, handing over at the end of a shift, making decisions with half the facts, and afterwards turning it into a review where nobody is blamed and something actually changes. It ends with the rules a regulated company (a bank, an insurer, a payment firm in the EU) has to follow on top of all that.
What you need to know already: 0.29 (running an incident), 0.30 and 0.31 (blameless postmortems), 0.16 (error budgets) and 29.15 (the first ten minutes). No new Linux: the work happens in the Incident console (Console in the left rail) or with the incident command in the terminal. Both drive the same room.
Hero mode, and why it fails
Hero mode is the default way small teams handle trouble: the most experienced person grabs the keyboard and fixes it, alone, fast, and tells everyone afterwards. It works often enough to feel right. It fails in four predictable ways:
- Nobody holds the whole picture. The hero is reading a stack trace. Nobody is watching whether the error rate moved, whether support knows, whether the change they are about to make collides with someone else's.
- Changes collide. Two well-meaning people restarting, scaling and rolling back at once produce a system nobody understands - and evidence nobody can trust.
- Silence. People outside the room hear nothing, so they guess, escalate, and interrupt the hero with "any update?", which slows the fix.
- It does not scale. The hero burns out, goes on holiday, or is the cause.
The answer is old, and it did not come from software.
Where incident command comes from
In the autumn of 1970 wildfires burned across Southern California for thirteen days: 16 people died, more than 700 buildings and over 500,000 acres burned. Dozens of fire agencies turned up, and the after-action reviews found that the fire was not the only problem. The agencies used different words for the same things, nobody knew who was in charge of whom, radios did not interoperate, and resources were sent twice or not at all.
The response was FIRESCOPE (Firefighting Resources of California Organized for Potential Emergencies), formed in 1972, which built the Incident Command System (ICS): one way to organise any emergency, whoever turns up. California agencies adopted it around 1980; the US made it national doctrine in the National Incident Management System (NIMS), first published in March 2004 and in its third edition since October 2017. Fire, police, hospitals and disaster relief all run on it today.
The ideas that carry straight over to a software incident (from the NIMS doctrine):
- Common terminology. Everyone uses the same words for roles and states. "SEV2", "IC", "ops lead", "mitigated", "resolved" mean one thing in your company.
- Modular organisation. The structure grows and shrinks with the incident. A small incident is one person wearing every hat; a big one splits them.
- Management by objectives. The commander sets specific, measurable objectives ("checkout errors under 1% by 20:30"), not a list of busy work.
- Unity of command. Every person reports to and takes direction from exactly one person. In practice: one IC, and only the ops lead's people touch production.
- Manageable span of control. NIMS gives the optimal span as one supervisor to five people (FEMA's training says three to seven). An IC with twelve people talking to them directly is not commanding, they are drowning: add a layer.
- Establishment and transfer of command. Someone explicitly takes command, and a handoff is a briefing plus telling everyone who is in charge now.
- Incident action planning in operational periods. The plan holds until a set time, then it is reviewed. Your version: "next update at 20:15".
Google's incident process is openly built on ICS - chapter 14 of the Site Reliability Engineering book, "Managing Incidents", says so - and so is PagerDuty's public Incident Response guide. The SRE Workbook sums up the job in three words, the 3Cs: coordinate the response, communicate inside and outside, keep control of what is being done.
The two jobs that must not share a brain
Chapter 0 named the roles. The single most important rule is the one people break first: the incident commander does not debug. PagerDuty's guide puts it bluntly - the IC is "NOT a resolver" and should not be "performing any actions or remediations, checking graphs, or investigating logs".
Why so strict? Debugging needs deep focus on one thing; commanding needs shallow attention on everything. The same brain cannot do both. The moment the IC opens a terminal, the questions "what is the impact now", "who is doing what", "when is the next update" have no owner. Google's book makes the same point the other way round: the operations team "should be the only group modifying the system during an incident" - so the IC does not, and neither does anyone else.
On a two-person team at 3am you will wear both hats. Then say which hat you have on ("IC hat: I am posting the update; ops hat: rolling back now") and take the IC hat back every few minutes to look at the whole.
Declaring is a decision, not a state
An incident does not start when something breaks. It starts when someone says "this is an incident, I am the IC". From that sentence on, the room has a commander, a severity, a cadence and a channel. Before it, it has four people debugging. Google's book lists three questions; if any answer is yes, declare:
Do you need to involve a second team in fixing the problem?
Is the outage visible to customers?
Is the issue unsolved even after an hour's concentrated analysis?
and adds: "It is better to declare an incident early and then find a simple fix and close out the incident than to have to spin up the incident management framework hours into a burgeoning problem." Declaring costs a message. Not declaring costs the first hour.
The incident room in this course
Real companies use a chat tool (a channel per incident), a paging tool (who is on call, who gets woken, escalation if they do not answer) and a status page (what customers see). The lab has all three in one place, the incident room:
- the Incident console (Console in the left rail, while a chapter 37 step is open): the channel, the incident card, buttons for every action and the status page as customers see it;
- the
incidentcommand in the terminal, for the same actions typed out (both are labelled(simulator); the practice is real, the tooling is a stand-in).
The room has its own clock. Incident time does not follow your wall clock: every action takes the minutes it would take in a real incident (posting a status update about four, paging one), and incident wait N lets N minutes pass. Responders reply when they would reply; pages are acknowledged late or not at all; support asks for news when the update you promised does not arrive.
Each lesson in this chapter opens a room of its own to try things in. This one holds INC-4519, the checkout failure from chapter 0's story, a minute after the page:
$ incident status
INC-4519 checkout failing after release 2.9.1
State: not declared clock 19:40 UTC (+0 min)
Alert: [FIRING] CheckoutErrorBudgetBurn severity=page
Impact now: 31.2% of checkout requests failing
Roles: IC - ops - comms - scribe -
In channel: @learner
Last update: none yet
Decisions: 0 notes: 0
Not declared. Assess the impact, then: incident declare --sev N "summary" (incident matrix shows the severity table)
The incident card, line by line: the incident number and title, whether it is declared and at what severity, the room's clock, the alert that started it, the impact the dashboard shows right now, who holds each role, who is in the channel, when each audience last heard from you, and the decisions and notes recorded so far.
The channel itself:
$ incident log
19:40 bot PagerDuty: [FIRING] CheckoutErrorBudgetBurn severity=page
19:40 bot checkout: 1h burn rate 312x the budget (31.2% of requests failing). Paged @learner (primary on-call, checkout).
bot lines are the tools talking (the pager, the incident bot); @name lines are people. The 312x is chapter 0's burn rate: a 31.2% error ratio against a 0.1% budget.
Who could you call?
$ incident oncall
TEAM LEVEL 1 LEVEL 2 LEVEL 3 ESCALATES AFTER WHAT
checkout @dana @mo @alex 5 min checkout and cart services
database @priya @omar @alex 15 min the orders and payments databases
platform @lee @casey @alex 5 min Kubernetes, CI, the shared platform
support @sam - - 10 min customer support (business hours + weekend rota)
PEOPLE
@dana Dana Okafor SRE, checkout team
@mo Mo Haddad developer, checkout team
@priya Priya Raman database on-call (primary)
@omar Omar Lindqvist database on-call (secondary)
@sam Sam Rivera support lead
@alex Alex Chen engineering manager
@lee Lee Novak SRE, platform team
Each team has an escalation policy: who is paged first (level 1), who next if level 1 does not acknowledge within the time on the right, and so on. Lesson 37.14 is about using it well.
You can do all of this in the console instead: the card is at the top right, the channel on the left, and every button shows the incident command it runs when you hover it. Use whichever you like; the labs check what happened in the room, not where you clicked.
incident help lists every subcommand; man incident has the details and examples.
In an interview
"Have you been incident commander?" is common for SRE and platform roles, and a junior answer that shows the structure beats a war story with no structure: "I declared early, took IC, gave the keyboard to one ops lead, kept the update cadence, and the postmortem fixed the pipeline, not the person." The questions behind it: why separate commanding from fixing, and why declare before you know the cause.
What you can now do
- explain hero mode and the four ways it fails
- say where incident command comes from (1970 California fires, FIRESCOPE, ICS, NIMS) and which of its ideas carry over: common terms, modular roles, unity of command, span of control, explicit transfer of command, objectives with a time
- explain why the IC does not debug, and use the three declare questions
- read the incident card, the channel and an escalation policy in the incident room