Why this lesson
You declared. You are the incident commander. Now what do your hands do, if they must not debug? This lesson is the IC's job minute by minute: the roles you hand out and to whom, how you give a task so it actually gets done, how you stop people from helping in ways that hurt, and the loop you run until it is over.
What you need to know already: 37.1 (why the IC does not debug), 37.2 (declaring), 0.29 (the four roles by name).
The roles, and their other names
Every guide splits the same work a little differently. The lab uses four roles, which map onto the others like this:
lab (and 0.29) Google SRE book PagerDuty guide Atlassian handbook
-------------- ------------------- -------------------------- -----------------------
IC Incident Command Incident Commander (+Deputy) Incident Manager
ops lead Operational Work Subject Matter Experts Tech Lead
comms lead Communication Customer + Internal Liaison Communications Manager
scribe Planning (part) Scribe (backup IM keeps timeline)
- Incident commander (IC). Owns the incident, not the fix. Sets and changes severity, hands out roles, decides (roll back or not, page more people or not, when it is over), keeps the cadence. Google: the commander "holds all positions that they have not delegated" - an empty role is the IC's job until someone takes it.
- Ops lead. The hands on production. Investigates and mitigates, with helpers if needed, and is the only route by which anything in production changes. Says what they are about to do before doing it.
- Comms lead. Writes and posts the updates to every audience on the cadence (37.8), answers support and stakeholders, and keeps them out of the responders' way. PagerDuty splits this into a customer liaison (the public status page) and an internal liaison (executives, legal, other teams).
- Scribe. Keeps the timeline: what was observed, decided and done, with times. Google calls the result the live incident state document and calls keeping it the IC's most important responsibility - which in practice means making sure someone does.
PagerDuty adds a deputy: a second trained IC who backs up the first, takes notes, and can take over. On a big incident it is the first role worth filling after ops.
Small incident, small team: one person holds all four. Second person arrives: split IC from ops first. Third: comms. Fourth: scribe. That order follows where the time goes - fixing, then talking, then writing down.
Giving a task that gets done
The bystander effect is real in incident channels: "can someone check the database?" in a room of eight people means nobody checks the database, or three people do. PagerDuty's IC training says never to ask "can someone"; name a person, say what you want, and say when you want to hear back:
weak: can someone look at the logs?
strong: @dana, look at the checkout errors and tell me what they say. Report back in 5 minutes.
Three parts: who, what (one thing, concrete), when you will check. The time box matters as much as the task: it tells the person how deep to go, and it tells you when silence means trouble.
Responders report back the same way. PagerDuty's guide has a name for the shape: a CAN report - Condition (what I see), Actions (what I did or am doing), Needs (what I need from you). "Errors are all TaxClient timeouts since 19:33. I have not changed anything. I need a decision on rolling back." Ask for it if people ramble.
When a decision is needed and the room is unsure, do not wait for everyone to agree. Make a proposal and poll for strong objections, the PagerDuty phrasing: "The proposal is to roll back to 2.9.0. Are there any strong objections?" A pause, then "hearing none, we proceed". Their guide is blunt about why: "Making the 'wrong' decision is better than making no decision."
One pair of hands on production
Unity of command (37.1) has a very concrete form: once there is an ops lead, nobody else changes production. Not "small" changes, not "just a restart". Two people changing the same system at once produce three problems: the changes interfere, the graphs stop meaning anything (did the rollback work, or the restart?), and the timeline becomes impossible to reconstruct.
If you join a room where several people are already changing things, the first useful act of command is to stop them, in plain words, to everyone at once: "Everyone: stop all changes to production. Nobody touches prod unless I assign it." Then hand the keyboard to one person. Incident 37.6 is exactly that room.
Running the room in the lab
You can only give a role to someone who is in the channel - which is how it works in real tools too: you page them first. In this lesson's room (INC-4519 again, fresh):
$ incident declare --sev 2 "checkout failing for about a third of customers since 19:33"
Declared INC-4519 SEV2 at 19:40 UTC.
19:40 bot INC-4519 declared SEV2 by @learner - you are incident commander until you hand it on. Channel #inc-4519-checkout. Policy: updates every 30 min (first due 20:10 UTC).
19:42 @sam Sam from support - I saw the incident come in. Customers are calling about failed payments. What can I tell them?
$ incident page checkout "checkout failing ~31% since 19:33, SEV2 declared"
Paged @dana.
19:42 bot PagerDuty: paged @dana (on call for checkout): "checkout failing ~31% since 19:33, SEV2 declared"
$ incident wait 2
19:43 -> 19:45 UTC
19:44 @dana Dana here, acked. What do you need from me?
Paging a team (checkout) pages whoever is on call for it right now - here @dana. Paging a person directly (incident page @dana) works too, but a team page follows the rota and its escalation policy (37.14). The message travels with the page: say what is wrong, not just "help".
$ incident role ops @priya
incident: @priya is not in the channel. Page them first: incident page @priya
$ incident role ops @dana
OPS: @dana
19:45 bot @dana is now ops lead (the only hands on production)
19:46 @dana Ops lead, got it. I am the only one changing production until you say otherwise. What first?
Now watch what happens when the IC reaches for the keyboard:
$ incident assign me "read the checkout logs"
Task #1 is yours. (You are IC: while your hands are on a keyboard, nobody is watching the whole incident.)
19:47 @dana If you are digging into that yourself, who is running the incident? Give me the hands-on part - you keep the big picture.
The room lets you (sometimes you must), and it says what it costs. The better move:
$ incident assign @dana "look at the checkout errors and tell me what they say - report back in 5 minutes"
Task #2 -> @dana
19:48 @dana On it, will report back.
$ incident wait 3
19:48 -> 19:51 UTC
19:49 bot Your task #1 (read the checkout logs): The errors are all one kind: "TaxClient: timeout after 2000ms" from the new tax service call. They start at 19:33, two minutes after 2.9.1 went out.
19:50 @dana The errors are all one kind: "TaxClient: timeout after 2000ms" from the new tax service call. They start at 19:33, two minutes after 2.9.1 went out.
Your own task and Dana's found the same thing: the work was done twice, and for those minutes nobody was watching the impact or the clock. Who, what, when. Dana's answer is a condition you can act on: one kind of error, starting two minutes after a release. That points at a rollback - the next lesson's mission is the whole incident, page to resolve.
$ incident status
INC-4519 checkout failing after release 2.9.1
State: open SEV2 clock 19:51 UTC (+11 min)
Alert: [FIRING] CheckoutErrorBudgetBurn severity=page
Summary: checkout failing for about a third of customers since 19:33
Impact now: 33.8% of checkout requests failing
Roles: IC @learner ops @dana comms - scribe -
In channel: @learner @sam @dana
Last update: none yet
Next update: due 20:10 UTC (in 19 min)
Decisions: 0 notes: 0
The IC's loop
PagerDuty trains its incident commanders on a four-step loop, run again and again until the incident is over:
size-up what is the impact now? what do we know? who is here?
stabilize is anyone changing things uncoordinated? is the most promising
mitigation assigned to one owner with a time box?
update has every audience heard from us within the cadence?
verify did the last action do what we expected? (look at the impact, not the code)
Each pass takes a few minutes. Between passes the IC listens, answers, and writes down decisions. The questions the IC keeps asking (from 29.15): what is the impact, what do we know, what are we trying, who is doing it, when do we next check in.
When someone more senior arrives
A director joins and starts giving instructions. Two IC commands now. PagerDuty's answer is a question, asked politely and in the open: "Do you wish to take command?" If yes, do a proper handoff (37.14); if no, they are a stakeholder and get the exec update like everyone else. Their guide is explicit that "the arrival of a more qualified person does NOT necessarily mean a change in incident command". Incident 37.11 puts the CEO in the channel.
In an interview
"What does an incident commander actually do if they do not fix anything?" - coordinate, communicate, control: set severity, assign one ops lead and other roles by name, give time-boxed tasks, stop uncoordinated changes, keep the update cadence, make decisions (polling for strong objections), and decide when it is over.
What you can now do
- map the lab's four roles onto Google's, PagerDuty's and Atlassian's names, and fill them in the right order as people arrive
- give a task with who, what and when, and ask for a CAN report back
- enforce one pair of hands on production
- run the size-up / stabilize / update / verify loop, and answer "do you wish to take command?" when someone senior arrives