Chapter 31 Incident Command & Communication
Take command of an incident and run it: severity by impact, roles, updates on a clock for every audience, paging and handoffs, decisions under uncertainty, blameless postmortems, honest metrics and EU DORA reporting.
In plain words
Think of a restaurant kitchen at the busiest moment of the night, when an order goes badly wrong. In a bad kitchen every cook drops what they are doing and grabs the same pan, the waiters tell each table something different, and nobody knows who decided what. In a good kitchen the head chef does not cook during the rush: they call the orders, one cook owns each station, a waiter tells the tables how long, and someone keeps the tickets.
An incident is that rush. This chapter is the people half of it: declaring with a severity judged by user impact, one incident commander who does not debug, one ops lead with hands on production, updates on a clock for every audience, paging and escalating, handing over at the end of a shift, deciding with half the facts, and a blameless review afterwards. You practise in the incident room, through the incident command or the Incident console.
Why it matters on call
The tools find the problem; how long users suffer depends on how the incident is run. In SRE and platform roles you will be paged, and sooner or later you will be the person who has to say "this is an incident, I am the IC". Knowing the structure (roles, cadence, escalation, handoffs, decisions on the record) is what turns a room of four people changing production at once into a response.
It is also a staple of interviews: "have you been incident commander?", "how do you decide severity?", "how would you measure incident response?". A junior answer that shows the structure beats a war story without one. And if you work for a bank, an insurer or a payment firm in the EU, the last lesson matters on its own: DORA adds a regulator's clock next to your update cadence, and the IC's timeline is where its timestamps come from.
Lessons
- Why incident command: from wildfires to the IC who does not debug
- Severity levels: the impact matrix, and declaring early
- Roles: the incident commander coordinates, does not debug
- Communication: cadence, audiences and templates
- Paging, escalation and handoffs
- Deciding under uncertainty: mitigate first, rollback bias, OODA
- Ending the incident, and the postmortem that changes something
- Incident metrics and their traps; on-call health
- Regulated environments: DORA major ICT incident reporting
24 hands-on labs (missions, incidents and drills) run in the terminal: Open this chapter in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.
Questions people ask
Do I need to be senior to be incident commander?
No. The IC owns the incident, not the fix: severity, roles, decisions and the cadence. Whoever declares is IC until they hand it on. A junior who declares early, gives the keyboard to one ops lead and keeps the updates going is doing the job well. When someone more senior arrives, ask "Do you wish to take command?"; PagerDuty's guide says the arrival of a more qualified person does not necessarily mean a change in command.
Is the incident room a real tool?
No, it is a stand-in, labelled (simulator). Real companies use a chat tool with a channel per incident, a paging tool with rotas and escalation, and a status page. The room puts all three in one place: the Incident console in the left rail, or the incident command in the terminal, which drive the same room. It has its own clock: every action takes the minutes it would take for real, and incident wait N lets time pass.
How does this chapter relate to chapter 0?
Chapter 0 (0.29 to 0.31) taught the words: severity, the incident commander, mitigate first, the update cadence and blameless postmortems; lesson 29.15 turned them into the first ten minutes and a template. Those lessons quietly assumed people do what the process says. This chapter is about the room where they do not: taking command of an incident already in motion, six audiences on a clock, escalation, handoffs and reviews that change something.
Why are there two things called DORA in this chapter?
Same four letters, unrelated things. The DORA metrics come from Google's DevOps Research and Assessment programme and measure software delivery; one of them is now called failed deployment recovery time. The other DORA is the EU's Digital Operational Resilience Act, Regulation (EU) 2022/2554, which has applied since 17 January 2025 and makes financial entities classify and report major ICT-related incidents to their supervisor.
Is all this process too much for a small team?
The structure is modular: it grows and shrinks with the incident, which is one of the ICS ideas the chapter borrows. A small incident is one person wearing every hat. When a second person arrives, split IC from ops first; with a third, add a comms lead; with a fourth, a scribe. What never shrinks is the habit: declare, say who is in charge, keep one pair of hands on production and promise the next update.