OnCallReady

Lesson 31.23 · Incident Command & Communication · 21 min read

Ending the incident, and the postmortem that changes something

In plain words

A football team loses a match and sits down to watch the replay. A bad coach looks for the player to bench. A good coach asks why there was a gap in the defence: was the formation wrong, did a signal get missed, was everyone tired after a cup game? If the players expect to be punished they will say as little as possible, and the same gap opens next week. If they can say "I thought the winger was covering, because that is what we practised", the team learns something.

A postmortem is that replay. First you end the incident properly: resolved only when the impact has ended, every audience told, responders thanked. Then a blameless review from the timeline: the second story instead of "human error", contributing factors instead of one root cause, and action items with an owner, a date and a way to tell they are done.

Why this lesson

The incident is over when users are fine. The work is over when the system is less likely to do it again - and that only happens if the end of the incident and the review afterwards are run as carefully as the response. Chapter 0 (0.30, 0.31) taught why postmortems are blameless and how to write one from evidence; 29.15 gave a template. This lesson is the part those skipped: ending the incident well, why "root cause" is the wrong question, how to run the review meeting, and how to write action items that actually ship.

What you need to know already: 0.30 and 0.31 (blameless postmortems, the timeline), 29.15 (the template), 37.20 (decisions on the record).

Ending the incident

Two different moments, often confused:

Close the incident only when the metrics have been healthy for a while (the lab asks for 15 minutes) and nothing is about to break again. Then close it properly - a short checklist the IC runs before leaving:

status page     final "Resolved" post: start, end, what customers saw, an apology;
                every component back to operational
every audience  internal, exec and support hear "resolved" too, from you
postmortem      an owner and a review date (in the room: incident postmortem)
follow-ups      the obvious tickets filed now, while everyone remembers
people          thank the responders by name; tell whoever was paged they can sleep

Google's book lists when a postmortem is required: user-visible downtime or degradation beyond a threshold; data loss of any kind; on-call intervention such as a release rollback or rerouting traffic; a resolution time above a threshold; and a monitoring failure (a human found it before an alert did). It adds that any stakeholder may request one. Atlassian requires one for every SEV1 and SEV2.

The timeline is the raw material

The postmortem starts from the timeline, not from memory. A scribe (or the room itself) has been recording it; in the lab, incident timeline prints what the room saw you do, and --md gives it as a markdown list to paste into the document. A short run of INC-4519:

$ incident declare --sev 2 "checkout failing for about a third of customers since 19:33"
Declared INC-4519 SEV2 at 19:40 UTC.
19:40  bot       INC-4519 declared SEV2 by @learner - you are incident commander until you hand it on. Channel #inc-4519-checkout. Policy: updates every 30 min (first due 20:10 UTC).
19:42  @sam      Sam from support - I saw the incident come in. Customers are calling about failed payments. What can I tell them?
$ incident page checkout
Paged @dana.
19:42  bot       PagerDuty: paged @dana (on call for checkout)
$ incident wait 2
19:43 -> 19:45 UTC
19:44  @dana     Dana here, acked. What do you need from me?
$ incident role ops @dana
OPS: @dana
19:45  bot       @dana is now ops lead (the only hands on production)
19:46  @dana     Ops lead, got it. I am the only one changing production until you say otherwise. What first?
$ incident decide "roll back checkout to 2.9.0" --why "failures started two minutes after 2.9.1"
Decision recorded at 19:46 UTC.
$ incident assign @dana "roll back checkout to 2.9.0"
Task #1 -> @dana
19:48  @dana     On it: roll back checkout to 2.9.0. I'll report back.
$ incident wait 9
19:48 -> 19:57 UTC
19:52  @alex     Alex here (eng manager). Leadership is asking about checkout - can I get two lines I can forward? Impact, what we are doing, when you will know more.
19:53  @dana     Rollback finished: checkout is on 2.9.0 on all 6 pods. Errors are dropping on the dashboard.
$ incident timeline --md
## Timeline (2026-09-29, UTC)
- 19:40  page: [FIRING] CheckoutErrorBudgetBurn severity=page
- 19:40  declared SEV2: checkout failing for about a third of customers since 19:33
- 19:42  paged @dana
- 19:44  @dana acknowledged
- 19:45  OPS: @dana
- 19:46  decision: roll back checkout to 2.9.0
- 19:47  task -> @dana: roll back checkout to 2.9.0
- 19:53  @dana done: Rollback finished: checkout is on 2.9.0 on all 6 pods. Errors are dropping on the dashboard.
- 19:55  dashboard: back under 1% (checkout requests failing)

What the room cannot record, people add: what they saw on which dashboard, what they expected an action to do, the moments of confusion. Those are often the most useful lines of all.

Blameless, one level deeper

You know the rule: the postmortem is about the system, never a person to blame. The reason it works is less obvious than "be nice".

In 2012 John Allspaw, then at Etsy, described the goal as a just culture: balancing safety and accountability by having engineers give "a detailed account of" what actions they took at what time, what effects they observed, the expectations and assumptions they had, and their understanding of the timeline - "without fear of punishment or retribution". People who expect to be punished leave things out. Then nobody learns why the action that broke production looked right at the time, and the next engineer does it again.

Two biases make blame feel fair when it is not:

Allspaw calls the alternative the second story: the first story is "human error caused the outage"; the second asks what about the system made that error easy, likely or invisible. Google's book says it shortest: "You can't 'fix' people, but you can fix systems and processes."

Blameless is not consequence-free. Accountability moves forward in time: to the action items, which have owners and dates and are tracked until they are done.

Root cause, and why the singular is wrong

A postmortem template usually has a "root cause" heading, and people fill it with one line. Richard Cook's point 7 is blunt: "Post-accident attribution to a 'root cause' is fundamentally wrong." Complex systems fail when several things line up - the release had a slow call, the call had no timeout budget, the pipeline had no canary, the alert fired but the rollback runbook was stale. Remove any one and the incident is smaller or does not happen.

So write contributing factors, plural, and keep asking until you reach ones you can change. Two tools, used with care:

weak:   Root cause: Mo deployed 2.9.1 without testing the tax service.
better: Contributing factors:
        - 2.9.1 added a synchronous call to the tax service with a 2 s timeout
        - the release went to 100% at once: checkout has no canary stage
        - the tax service's staging copy answers in 20 ms, production in 1.8-2.5 s
        - the rollback runbook referenced a command removed in July

Running the review meeting

The document is drafted before the meeting, by its owner, from the timeline. The meeting is for understanding and for agreeing the actions - usually an hour, within a week of the incident. Someone facilitates who was not the IC (the IC is a participant with a lot to say). A workable agenda:

5 min    purpose, ground rules: we are here to understand the system; nobody is on trial
20 min   walk the timeline; at each decision: what did you know, what did you expect?
15 min   contributing factors, what went well, where we got lucky
15 min   action items: owner, date, type, ticket - agreed in the room
5 min    who sends the summary, when the actions are reviewed again

The facilitator's main tool is the question. "Why did you deploy on a Friday?" puts a person on the defensive. "What made deploying then look like the right call?" gets the second story: the pipeline was green, the change was small, the team deploys daily. Ask "how" and "what" more than "why you". Refer to people by role ("the on-call", "the release owner") in the document - Atlassian's practice too.

When someone asks "who broke it?" - and in a real meeting someone will, often a manager under pressure - answer the need behind the question: they want to know it will not happen again. "It is not about who; the question is how one config change could reach every customer without a canary. That is what we are fixing, and the owner of that fix is ..." Mission 37.24 puts you in exactly that conversation.

Action items that ship

Most postmortems fail here: twenty items like "improve monitoring" that nobody owns and nobody closes. Atlassian splits follow-ups into priority actions, which carry a deadline (an internal target of 4 or 8 weeks, depending on the service), and the rest. Google's Lueder and Beyer ("Postmortem Action Items: Plan the Work and Work the Plan") want each item actionable (it starts with a verb and says what to do), specific and bounded (it is clear when it is done). The common management acronym says the same thing: SMART - specific, measurable, achievable, relevant, time-bound.

The lab's checker wants five things on every item, plus no non-actions:

owner       @team or @person - one name, not "the team"
due         a date, YYYY-MM-DD
done when   something you can check: a number, a test, an alert that fires
type        prevent (it cannot happen again), detect (we see it sooner),
            mitigate (it hurts less / is shorter), process (how we work)
ticket      the tracking id, so it is followed up like any other work
$ incident check action "Improve monitoring of checkout"
action items (simulator - owner, due date, done-when, type, ticket; no non-actions):
FIX     Improve monitoring of checkout
         - no owner (@team or @person)
         - no due date (YYYY-MM-DD)
         - no way to tell it is done (a number, a test, "done when ...")
         - no type (prevent / detect / mitigate / process)
         - no ticket id
         - a non-action ("be more careful", "improve monitoring")
Not actionable yet.
$ incident check action "Add a canary stage to the checkout pipeline that blocks the release when the canary's error rate is over 1% for 5 min - owner: @dana - due 2026-10-20 - type: prevent - CHK-2291"
action items (simulator - owner, due date, done-when, type, ticket; no non-actions):
ok      Add a canary stage to the checkout pipeline that blocks the release when the canary's error rate is over 1% for 5 min - owner: @dana - due 2026-10-20 - type: prevent - CHK-2291
All actionable.

What you can now do

Why it helps

The incident is over when users are fine; the work is over when the system is less likely to do it again. That only happens if the review is run as carefully as the response. In interviews "how do you run a blameless postmortem?" is common, and a manager asking "who broke it?" in the review meeting is something you will actually face.

This lesson gives you the vocabulary and the moves: just culture and why people who expect punishment leave things out, hindsight bias and the fundamental attribution error, Richard Cook's case against a single root cause, how to facilitate with "what made that look like the right call?" instead of "why did you?", and how to turn "improve monitoring" into an action item that ships, with an owner, a due date, a done-when, a type and a ticket.

Commands in this lesson

incident

FAQ

What is the difference between mitigated and resolved?

Mitigated means users are no longer hurt, but the underlying problem may still be there: the bad release is rolled back, the bug in it is not fixed. Resolved, in Atlassian's words, is when the current or imminent business impact has ended. Not when the fix is merged and not when the graph first turns green: the lab asks for 15 healthy minutes. Then close properly: a final status post, every audience, a postmortem owner, follow-ups, thanks.

When is a postmortem required?

Google's book lists the triggers: user-visible downtime or degradation beyond a threshold, data loss of any kind, on-call intervention such as a release rollback or rerouting traffic, a resolution time above a threshold, and a monitoring failure, where a human found it before an alert did. Any stakeholder may also request one. Atlassian requires a postmortem for every SEV1 and SEV2.

Does blameless mean nobody is accountable?

No. John Allspaw described the goal in 2012 as a just culture: engineers give a detailed account of what they did, what they saw, what they expected and assumed, without fear of punishment, because people who expect to be punished leave things out and nobody learns why the action looked right at the time. Accountability moves forward in time, to the action items, which have owners and dates and are tracked until done.

Why is "root cause" the wrong question?

Richard Cook's How Complex Systems Fail: post-accident attribution to a root cause is fundamentally wrong. Complex systems fail when several things line up: a slow call, no timeout budget, no canary, a stale rollback runbook. Write contributing factors, plural. Five whys can help but follows one chain and stops at the first satisfying answer, so ask it down several branches. Atlassian's root cause is about where to act, not who is guilty.

What makes an action item one that actually ships?

Google's Lueder and Beyer want each item actionable (starts with a verb), specific and bounded (clear when it is done); SMART says the same. The lab's checker wants an owner (one name), a due date, a done-when you can check, a type (prevent, detect, mitigate, process) and a ticket. "Be more careful when deploying" and "improve monitoring" are non-actions. Atlassian gives priority actions a deadline of 4 or 8 weeks.

In an interview Mid

How do you run a blameless postmortem that actually changes something?

The owner drafts it before the meeting from the timeline, not from memory. Someone who was not the IC facilitates, within a week, about an hour: ground rules, walk the timeline asking at each decision what people knew and expected, contributing factors and what went well, then action items agreed in the room.

Blameless because people who expect punishment leave things out. I look for the second story - what made the action easy, likely or invisible - and watch for hindsight bias and the fundamental attribution error. I ask "what made deploying then look like the right call?", never "why did you deploy on a Friday?", and refer to people by role. If a manager asks who broke it, I answer the need: it will not happen again, and here is the owner of the fix.

No single root cause: contributing factors, plural, until we reach ones we can change. Accountability moves to the action items: each with one owner, a due date, a done-when, a type (prevent, detect, mitigate, process) and a ticket, tracked until done. "Improve monitoring" does not count.

Also asked: What do you say when a manager asks who broke production? · How do you make sure postmortem action items actually get done? · What is the difference between mitigated and resolved?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.