OnCallReady

Lesson 31.4 · Incident Command & Communication · 21 min read

Roles: the incident commander coordinates, does not debug

In plain words

Watch a busy kitchen during the dinner rush. The head chef stands at the pass: calls the orders, watches every plate, decides what goes out. They do not cook. Each cook owns one station, and nobody reaches into someone else's pan. A waiter tells the tables how long. Someone keeps the tickets in order. If the chef shouts "can someone get the sauce?", nobody does, or three people do; "Ana, sauce for table four, two minutes" gets one sauce.

Incident roles are that kitchen. The incident commander owns the incident, not the fix. The ops lead is the only pair of hands on production. The comms lead talks to every audience on the cadence. The scribe keeps the timeline. The IC gives tasks with who, what and when, asks for a CAN report back, polls for strong objections before a decision, and runs a loop: size-up, stabilize, update, verify.

Why this lesson

You declared. You are the incident commander. Now what do your hands do, if they must not debug? This lesson is the IC's job minute by minute: the roles you hand out and to whom, how you give a task so it actually gets done, how you stop people from helping in ways that hurt, and the loop you run until it is over.

What you need to know already: 37.1 (why the IC does not debug), 37.2 (declaring), 0.29 (the four roles by name).

The roles, and their other names

Every guide splits the same work a little differently. The lab uses four roles, which map onto the others like this:

lab (and 0.29)   Google SRE book       PagerDuty guide              Atlassian handbook
--------------   -------------------   --------------------------   -----------------------
IC               Incident Command      Incident Commander (+Deputy) Incident Manager
ops lead         Operational Work      Subject Matter Experts       Tech Lead
comms lead       Communication         Customer + Internal Liaison  Communications Manager
scribe           Planning (part)       Scribe                       (backup IM keeps timeline)

PagerDuty adds a deputy: a second trained IC who backs up the first, takes notes, and can take over. On a big incident it is the first role worth filling after ops.

Small incident, small team: one person holds all four. Second person arrives: split IC from ops first. Third: comms. Fourth: scribe. That order follows where the time goes - fixing, then talking, then writing down.

Giving a task that gets done

The bystander effect is real in incident channels: "can someone check the database?" in a room of eight people means nobody checks the database, or three people do. PagerDuty's IC training says never to ask "can someone"; name a person, say what you want, and say when you want to hear back:

weak:    can someone look at the logs?
strong:  @dana, look at the checkout errors and tell me what they say. Report back in 5 minutes.

Three parts: who, what (one thing, concrete), when you will check. The time box matters as much as the task: it tells the person how deep to go, and it tells you when silence means trouble.

Responders report back the same way. PagerDuty's guide has a name for the shape: a CAN report - Condition (what I see), Actions (what I did or am doing), Needs (what I need from you). "Errors are all TaxClient timeouts since 19:33. I have not changed anything. I need a decision on rolling back." Ask for it if people ramble.

When a decision is needed and the room is unsure, do not wait for everyone to agree. Make a proposal and poll for strong objections, the PagerDuty phrasing: "The proposal is to roll back to 2.9.0. Are there any strong objections?" A pause, then "hearing none, we proceed". Their guide is blunt about why: "Making the 'wrong' decision is better than making no decision."

One pair of hands on production

Unity of command (37.1) has a very concrete form: once there is an ops lead, nobody else changes production. Not "small" changes, not "just a restart". Two people changing the same system at once produce three problems: the changes interfere, the graphs stop meaning anything (did the rollback work, or the restart?), and the timeline becomes impossible to reconstruct.

If you join a room where several people are already changing things, the first useful act of command is to stop them, in plain words, to everyone at once: "Everyone: stop all changes to production. Nobody touches prod unless I assign it." Then hand the keyboard to one person. Incident 37.6 is exactly that room.

Running the room in the lab

You can only give a role to someone who is in the channel - which is how it works in real tools too: you page them first. In this lesson's room (INC-4519 again, fresh):

$ incident declare --sev 2 "checkout failing for about a third of customers since 19:33"
Declared INC-4519 SEV2 at 19:40 UTC.
19:40  bot       INC-4519 declared SEV2 by @learner - you are incident commander until you hand it on. Channel #inc-4519-checkout. Policy: updates every 30 min (first due 20:10 UTC).
19:42  @sam      Sam from support - I saw the incident come in. Customers are calling about failed payments. What can I tell them?
$ incident page checkout "checkout failing ~31% since 19:33, SEV2 declared"
Paged @dana.
19:42  bot       PagerDuty: paged @dana (on call for checkout): "checkout failing ~31% since 19:33, SEV2 declared"
$ incident wait 2
19:43 -> 19:45 UTC
19:44  @dana     Dana here, acked. What do you need from me?

Paging a team (checkout) pages whoever is on call for it right now - here @dana. Paging a person directly (incident page @dana) works too, but a team page follows the rota and its escalation policy (37.14). The message travels with the page: say what is wrong, not just "help".

$ incident role ops @priya
incident: @priya is not in the channel. Page them first: incident page @priya
$ incident role ops @dana
OPS: @dana
19:45  bot       @dana is now ops lead (the only hands on production)
19:46  @dana     Ops lead, got it. I am the only one changing production until you say otherwise. What first?

Now watch what happens when the IC reaches for the keyboard:

$ incident assign me "read the checkout logs"
Task #1 is yours. (You are IC: while your hands are on a keyboard, nobody is watching the whole incident.)
19:47  @dana     If you are digging into that yourself, who is running the incident? Give me the hands-on part - you keep the big picture.

The room lets you (sometimes you must), and it says what it costs. The better move:

$ incident assign @dana "look at the checkout errors and tell me what they say - report back in 5 minutes"
Task #2 -> @dana
19:48  @dana     On it, will report back.
$ incident wait 3
19:48 -> 19:51 UTC
19:49  bot       Your task #1 (read the checkout logs): The errors are all one kind: "TaxClient: timeout after 2000ms" from the new tax service call. They start at 19:33, two minutes after 2.9.1 went out.
19:50  @dana     The errors are all one kind: "TaxClient: timeout after 2000ms" from the new tax service call. They start at 19:33, two minutes after 2.9.1 went out.

Your own task and Dana's found the same thing: the work was done twice, and for those minutes nobody was watching the impact or the clock. Who, what, when. Dana's answer is a condition you can act on: one kind of error, starting two minutes after a release. That points at a rollback - the next lesson's mission is the whole incident, page to resolve.

$ incident status
INC-4519  checkout failing after release 2.9.1
State:        open  SEV2   clock 19:51 UTC (+11 min)
Alert:        [FIRING] CheckoutErrorBudgetBurn severity=page
Summary:      checkout failing for about a third of customers since 19:33
Impact now:   33.8% of checkout requests failing
Roles:        IC @learner   ops @dana   comms -   scribe -
In channel:   @learner @sam @dana
Last update:  none yet
Next update:  due 20:10 UTC (in 19 min)
Decisions:    0   notes: 0

The IC's loop

PagerDuty trains its incident commanders on a four-step loop, run again and again until the incident is over:

size-up     what is the impact now? what do we know? who is here?
stabilize   is anyone changing things uncoordinated? is the most promising
            mitigation assigned to one owner with a time box?
update      has every audience heard from us within the cadence?
verify      did the last action do what we expected? (look at the impact, not the code)

Each pass takes a few minutes. Between passes the IC listens, answers, and writes down decisions. The questions the IC keeps asking (from 29.15): what is the impact, what do we know, what are we trying, who is doing it, when do we next check in.

When someone more senior arrives

A director joins and starts giving instructions. Two IC commands now. PagerDuty's answer is a question, asked politely and in the open: "Do you wish to take command?" If yes, do a proper handoff (37.14); if no, they are a stakeholder and get the exec update like everyone else. Their guide is explicit that "the arrival of a more qualified person does NOT necessarily mean a change in incident command". Incident 37.11 puts the CEO in the channel.

In an interview

"What does an incident commander actually do if they do not fix anything?" - coordinate, communicate, control: set severity, assign one ops lead and other roles by name, give time-boxed tasks, stop uncoordinated changes, keep the update cadence, make decisions (polling for strong objections), and decide when it is over.

What you can now do

Why it helps

"What does an incident commander actually do if they do not fix anything?" is a real interview question, and this lesson is the answer. It also gives you the moves that make a room work: naming a person instead of "can someone" (the bystander effect is real in incident channels), a time box on every task so silence means something, and a clear stop when several people are changing production.

You will meet different names for the same roles: Google's Incident Command and Operational Work, PagerDuty's Subject Matter Experts and liaisons, Atlassian's Incident Manager and Tech Lead. Knowing the mapping lets you work with any company's process on day one. And the polite question for a senior who starts giving orders, "Do you wish to take command?", saves you from having two incident commanders in one room.

Commands in this lesson

incident

FAQ

Which role do I fill first when people arrive?

Split IC from ops first: the second person takes the hands-on work so the IC can watch the whole. The third becomes comms lead, the fourth scribe. That order follows where the time goes: fixing, then talking, then writing down. Until a role is filled it is the IC's job; Google's book says the commander holds all positions they have not delegated. On a big incident, PagerDuty's deputy is the first role worth adding after ops.

What is a CAN report?

PagerDuty's shape for reporting back: Condition (what I see), Actions (what I did or am doing), Needs (what I need from you). For example: "Errors are all TaxClient timeouts since 19:33. I have not changed anything. I need a decision on rolling back." It is short, it separates facts from actions, and it ends with something the IC can act on. Ask for it when people ramble.

Why should I never ask "can someone look at this"?

Because of the bystander effect: in a room of eight people, "can someone check the database?" means nobody checks it, or three people do. Give every task three parts: who (a name), what (one concrete thing) and when you will check. The time box tells the person how deep to go and tells you when silence means trouble: "@dana, look at the checkout errors and tell me what they say. Report back in 5 minutes."

What does polling for strong objections mean?

A way to decide without waiting for everyone to agree. State a proposal and ask "Are there any strong objections?", pause, then "hearing none, we proceed". It gives people a real chance to stop a bad idea without turning the call into a debate. PagerDuty's guide explains why: "Making the 'wrong' decision is better than making no decision." In the lab you can record the decision first with incident decide.

What do I do when a director joins and starts giving instructions?

Ask, politely and in the open: "Do you wish to take command?" If yes, do a proper handoff with a briefing and announce who is in charge now. If no, they are a stakeholder and get the executive update like everyone else, while instructions keep coming from one IC. PagerDuty's guide is explicit that the arrival of a more qualified person does not necessarily mean a change in incident command.

In an interview Junior

What does an incident commander actually do if they do not fix anything?

Coordinate, communicate, control. Concretely:

Between those I run a loop: size-up (impact now, who is here), stabilize (anyone changing things uncoordinated, the best mitigation owned and time-boxed), update (has everyone heard within the cadence), verify (did the last action do what we expected).

Also asked: How do you stop several people changing production at once? · What is the difference between the ops lead and the incident commander? · What do you do when someone more senior joins the incident?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.