OnCallReady

Lesson 31.1 · Incident Command & Communication · 18 min read

Why incident command: from wildfires to the IC who does not debug

In plain words

Picture a car breaking down with four friends inside. Everyone jumps out at once. One opens the bonnet, one phones a tow truck, another phones a different tow truck, and someone tells the family "we will be there in ten minutes" without asking anyone. Nobody knows what the others did, so two trucks arrive and the family is still waiting at midnight. Now imagine one friend says: "I will run this. You look at the engine, you call one truck, I will text the family at six." Same car, same problem, much faster.

That is the difference between hero mode and incident command. The idea comes from firefighters: after the 1970 California wildfires, FIRESCOPE built the Incident Command System so any number of agencies could work under one structure. Software borrowed it: one incident commander who coordinates and does not debug, one person changing production, and a declared incident with a cadence.

Why this chapter exists

You already know the words. Chapter 0 (0.29) taught severity, the incident commander, mitigating first and the update cadence; lesson 29.15 turned them into the first ten minutes of an incident and a postmortem template. Every one of those lessons had a quiet assumption: that the people involved do what the process says.

They do not. At 18:55 on a Tuesday four engineers are changing production at the same time, each convinced they are helping. The support lead is asking "is something going on?" in three channels. A manager wants a one-liner for the executives. Nobody has declared anything, so nobody is in charge, so nobody can say "stop". The bug in that incident takes ten minutes to fix. The incident lasts an hour.

This chapter is about the people half of an incident: taking command of a room that is already in motion, saying the right thing to six audiences on a clock, paging the right person and escalating when they do not answer, handing over at the end of a shift, making decisions with half the facts, and afterwards turning it into a review where nobody is blamed and something actually changes. It ends with the rules a regulated company (a bank, an insurer, a payment firm in the EU) has to follow on top of all that.

What you need to know already: 0.29 (running an incident), 0.30 and 0.31 (blameless postmortems), 0.16 (error budgets) and 29.15 (the first ten minutes). No new Linux: the work happens in the Incident console (Console in the left rail) or with the incident command in the terminal. Both drive the same room.

Hero mode, and why it fails

Hero mode is the default way small teams handle trouble: the most experienced person grabs the keyboard and fixes it, alone, fast, and tells everyone afterwards. It works often enough to feel right. It fails in four predictable ways:

The answer is old, and it did not come from software.

Where incident command comes from

In the autumn of 1970 wildfires burned across Southern California for thirteen days: 16 people died, more than 700 buildings and over 500,000 acres burned. Dozens of fire agencies turned up, and the after-action reviews found that the fire was not the only problem. The agencies used different words for the same things, nobody knew who was in charge of whom, radios did not interoperate, and resources were sent twice or not at all.

The response was FIRESCOPE (Firefighting Resources of California Organized for Potential Emergencies), formed in 1972, which built the Incident Command System (ICS): one way to organise any emergency, whoever turns up. California agencies adopted it around 1980; the US made it national doctrine in the National Incident Management System (NIMS), first published in March 2004 and in its third edition since October 2017. Fire, police, hospitals and disaster relief all run on it today.

The ideas that carry straight over to a software incident (from the NIMS doctrine):

Google's incident process is openly built on ICS - chapter 14 of the Site Reliability Engineering book, "Managing Incidents", says so - and so is PagerDuty's public Incident Response guide. The SRE Workbook sums up the job in three words, the 3Cs: coordinate the response, communicate inside and outside, keep control of what is being done.

The two jobs that must not share a brain

Chapter 0 named the roles. The single most important rule is the one people break first: the incident commander does not debug. PagerDuty's guide puts it bluntly - the IC is "NOT a resolver" and should not be "performing any actions or remediations, checking graphs, or investigating logs".

Why so strict? Debugging needs deep focus on one thing; commanding needs shallow attention on everything. The same brain cannot do both. The moment the IC opens a terminal, the questions "what is the impact now", "who is doing what", "when is the next update" have no owner. Google's book makes the same point the other way round: the operations team "should be the only group modifying the system during an incident" - so the IC does not, and neither does anyone else.

On a two-person team at 3am you will wear both hats. Then say which hat you have on ("IC hat: I am posting the update; ops hat: rolling back now") and take the IC hat back every few minutes to look at the whole.

Declaring is a decision, not a state

An incident does not start when something breaks. It starts when someone says "this is an incident, I am the IC". From that sentence on, the room has a commander, a severity, a cadence and a channel. Before it, it has four people debugging. Google's book lists three questions; if any answer is yes, declare:

Do you need to involve a second team in fixing the problem?
Is the outage visible to customers?
Is the issue unsolved even after an hour's concentrated analysis?

and adds: "It is better to declare an incident early and then find a simple fix and close out the incident than to have to spin up the incident management framework hours into a burgeoning problem." Declaring costs a message. Not declaring costs the first hour.

The incident room in this course

Real companies use a chat tool (a channel per incident), a paging tool (who is on call, who gets woken, escalation if they do not answer) and a status page (what customers see). The lab has all three in one place, the incident room:

The room has its own clock. Incident time does not follow your wall clock: every action takes the minutes it would take in a real incident (posting a status update about four, paging one), and incident wait N lets N minutes pass. Responders reply when they would reply; pages are acknowledged late or not at all; support asks for news when the update you promised does not arrive.

Each lesson in this chapter opens a room of its own to try things in. This one holds INC-4519, the checkout failure from chapter 0's story, a minute after the page:

$ incident status
INC-4519  checkout failing after release 2.9.1
State:        not declared   clock 19:40 UTC (+0 min)
Alert:        [FIRING] CheckoutErrorBudgetBurn severity=page
Impact now:   31.2% of checkout requests failing
Roles:        IC -   ops -   comms -   scribe -
In channel:   @learner
Last update:  none yet
Decisions:    0   notes: 0

Not declared. Assess the impact, then: incident declare --sev N "summary"   (incident matrix shows the severity table)

The incident card, line by line: the incident number and title, whether it is declared and at what severity, the room's clock, the alert that started it, the impact the dashboard shows right now, who holds each role, who is in the channel, when each audience last heard from you, and the decisions and notes recorded so far.

The channel itself:

$ incident log
19:40  bot       PagerDuty: [FIRING] CheckoutErrorBudgetBurn severity=page
19:40  bot       checkout: 1h burn rate 312x the budget (31.2% of requests failing). Paged @learner (primary on-call, checkout).

bot lines are the tools talking (the pager, the incident bot); @name lines are people. The 312x is chapter 0's burn rate: a 31.2% error ratio against a 0.1% budget.

Who could you call?

$ incident oncall
TEAM          LEVEL 1      LEVEL 2      LEVEL 3      ESCALATES AFTER   WHAT
checkout      @dana        @mo          @alex        5 min             checkout and cart services
database      @priya       @omar        @alex        15 min            the orders and payments databases
platform      @lee         @casey       @alex        5 min             Kubernetes, CI, the shared platform
support       @sam         -            -            10 min            customer support (business hours + weekend rota)

PEOPLE
  @dana     Dana Okafor      SRE, checkout team
  @mo       Mo Haddad        developer, checkout team
  @priya    Priya Raman      database on-call (primary)
  @omar     Omar Lindqvist   database on-call (secondary)
  @sam      Sam Rivera       support lead
  @alex     Alex Chen        engineering manager
  @lee      Lee Novak        SRE, platform team

Each team has an escalation policy: who is paged first (level 1), who next if level 1 does not acknowledge within the time on the right, and so on. Lesson 37.14 is about using it well.

You can do all of this in the console instead: the card is at the top right, the channel on the left, and every button shows the incident command it runs when you hover it. Use whichever you like; the labs check what happened in the room, not where you clicked.

incident help lists every subcommand; man incident has the details and examples.

In an interview

"Have you been incident commander?" is common for SRE and platform roles, and a junior answer that shows the structure beats a war story with no structure: "I declared early, took IC, gave the keyboard to one ops lead, kept the update cadence, and the postmortem fixed the pipeline, not the person." The questions behind it: why separate commanding from fixing, and why declare before you know the cause.

What you can now do

Why it helps

This lesson gives you the reason behind every rule that follows. When you join a room where four engineers are changing production and nobody has declared, you need to know why that is the problem, not the bug: nobody holds the whole picture, changes collide, people outside hear nothing, and the hero does not scale.

The ICS ideas carry straight over and make good interview material: common terminology, modular organisation, management by objectives, unity of command, a span of control of about five, and an explicit transfer of command. Google's three declare questions give you a test you can apply in seconds: a second team needed, customers seeing it, or an hour without an answer. And "the incident commander does not debug" is the one rule people break first, so be able to defend it.

Commands in this lesson

incident

FAQ

What is hero mode, and why is it bad if it usually works?

Hero mode is the most experienced person grabbing the keyboard and fixing it alone, then telling everyone afterwards. It works often enough to feel right, and fails in four predictable ways: nobody holds the whole picture, several people's changes collide and spoil the evidence, everyone outside hears nothing and interrupts with "any update?", and it does not scale because the hero burns out, goes on holiday, or is the cause.

Where does incident command come from?

From firefighting. In autumn 1970 wildfires burned across Southern California for thirteen days; dozens of agencies used different words, radios did not interoperate and nobody knew who was in charge of whom. FIRESCOPE, formed in 1972, built the Incident Command System. California adopted it around 1980, and the US made it national doctrine in NIMS, first published in March 2004, third edition since October 2017.

What is span of control?

How many people one supervisor can direct well. NIMS gives the optimal span as one supervisor to five people, and FEMA's training says three to seven. In an incident it means an IC with twelve people talking to them directly is not commanding but drowning. The fix is a layer: an ops lead who directs the helpers working on production, and a comms lead who handles the stakeholders.

When should I declare an incident?

Google's book lists three questions, and if any answer is yes, declare: do you need a second team to fix it, is the outage visible to customers, is it still unsolved after an hour's concentrated analysis? It adds that it is better to declare early and close it after a simple fix than to spin up the process hours into a growing problem. Declaring costs a message; not declaring costs the first hour.

What if only two of us are awake at 3am?

Then you wear both hats, and that is fine. Say which hat you have on, out loud in the channel: "IC hat: I am posting the update; ops hat: rolling back now." Every few minutes put the IC hat back on and look at the whole: what is the impact now, is anyone else changing things, when is the next update due. When a third person joins, hand the ops work to them.

In an interview Junior

Why should the incident commander not debug?

Debugging needs deep focus on one thing; commanding needs shallow attention on everything, and the same brain cannot do both. The moment the IC opens a terminal, "what is the impact now", "who is doing what" and "when is the next update" have no owner. That is hero mode again, just with a title.

PagerDuty's guide says the IC is "NOT a resolver" and should not be checking graphs or investigating logs. Google's book says the operations team should be the only group modifying the system during an incident - unity of command: one IC, and only the ops lead's people touch production. The IC's job is the 3Cs: coordinate, communicate, control.

On a tiny team I would wear both hats, but say which one I have on and take the IC hat back every few minutes. As soon as a second person arrives, IC and ops are the first roles to split.

Also asked: Have you ever been incident commander, and what did you do? · When would you declare an incident instead of just fixing it quietly? · What goes wrong when several engineers change production at the same time?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.