Why this lesson
The incident is over when users are fine. The work is over when the system is less likely to do it again - and that only happens if the end of the incident and the review afterwards are run as carefully as the response. Chapter 0 (0.30, 0.31) taught why postmortems are blameless and how to write one from evidence; 29.15 gave a template. This lesson is the part those skipped: ending the incident well, why "root cause" is the wrong question, how to run the review meeting, and how to write action items that actually ship.
What you need to know already: 0.30 and 0.31 (blameless postmortems, the timeline), 29.15 (the template), 37.20 (decisions on the record).
Ending the incident
Two different moments, often confused:
- Mitigated: users are no longer hurt, but the underlying problem may still be there (the bad release is rolled back; the bug in it is not fixed).
- Resolved: Atlassian's handbook defines it well - "an incident is resolved when the current or imminent business impact has ended". Not when the fix is merged; not when the graph first turns green.
Close the incident only when the metrics have been healthy for a while (the lab asks for 15 minutes) and nothing is about to break again. Then close it properly - a short checklist the IC runs before leaving:
status page final "Resolved" post: start, end, what customers saw, an apology;
every component back to operational
every audience internal, exec and support hear "resolved" too, from you
postmortem an owner and a review date (in the room: incident postmortem)
follow-ups the obvious tickets filed now, while everyone remembers
people thank the responders by name; tell whoever was paged they can sleep
Google's book lists when a postmortem is required: user-visible downtime or degradation beyond a threshold; data loss of any kind; on-call intervention such as a release rollback or rerouting traffic; a resolution time above a threshold; and a monitoring failure (a human found it before an alert did). It adds that any stakeholder may request one. Atlassian requires one for every SEV1 and SEV2.
The timeline is the raw material
The postmortem starts from the timeline, not from memory. A scribe (or the room itself) has been recording it; in the lab, incident timeline prints what the room saw you do, and --md gives it as a markdown list to paste into the document. A short run of INC-4519:
$ incident declare --sev 2 "checkout failing for about a third of customers since 19:33"
Declared INC-4519 SEV2 at 19:40 UTC.
19:40 bot INC-4519 declared SEV2 by @learner - you are incident commander until you hand it on. Channel #inc-4519-checkout. Policy: updates every 30 min (first due 20:10 UTC).
19:42 @sam Sam from support - I saw the incident come in. Customers are calling about failed payments. What can I tell them?
$ incident page checkout
Paged @dana.
19:42 bot PagerDuty: paged @dana (on call for checkout)
$ incident wait 2
19:43 -> 19:45 UTC
19:44 @dana Dana here, acked. What do you need from me?
$ incident role ops @dana
OPS: @dana
19:45 bot @dana is now ops lead (the only hands on production)
19:46 @dana Ops lead, got it. I am the only one changing production until you say otherwise. What first?
$ incident decide "roll back checkout to 2.9.0" --why "failures started two minutes after 2.9.1"
Decision recorded at 19:46 UTC.
$ incident assign @dana "roll back checkout to 2.9.0"
Task #1 -> @dana
19:48 @dana On it: roll back checkout to 2.9.0. I'll report back.
$ incident wait 9
19:48 -> 19:57 UTC
19:52 @alex Alex here (eng manager). Leadership is asking about checkout - can I get two lines I can forward? Impact, what we are doing, when you will know more.
19:53 @dana Rollback finished: checkout is on 2.9.0 on all 6 pods. Errors are dropping on the dashboard.
$ incident timeline --md
## Timeline (2026-09-29, UTC)
- 19:40 page: [FIRING] CheckoutErrorBudgetBurn severity=page
- 19:40 declared SEV2: checkout failing for about a third of customers since 19:33
- 19:42 paged @dana
- 19:44 @dana acknowledged
- 19:45 OPS: @dana
- 19:46 decision: roll back checkout to 2.9.0
- 19:47 task -> @dana: roll back checkout to 2.9.0
- 19:53 @dana done: Rollback finished: checkout is on 2.9.0 on all 6 pods. Errors are dropping on the dashboard.
- 19:55 dashboard: back under 1% (checkout requests failing)
What the room cannot record, people add: what they saw on which dashboard, what they expected an action to do, the moments of confusion. Those are often the most useful lines of all.
Blameless, one level deeper
You know the rule: the postmortem is about the system, never a person to blame. The reason it works is less obvious than "be nice".
In 2012 John Allspaw, then at Etsy, described the goal as a just culture: balancing safety and accountability by having engineers give "a detailed account of" what actions they took at what time, what effects they observed, the expectations and assumptions they had, and their understanding of the timeline - "without fear of punishment or retribution". People who expect to be punished leave things out. Then nobody learns why the action that broke production looked right at the time, and the next engineer does it again.
Two biases make blame feel fair when it is not:
- Hindsight bias. After the fact, the signs look obvious. At 19:31 with a green pipeline, they were not. (Cook's How Complex Systems Fail, point 8: "Hindsight biases post-accident assessments of human performance.")
- The fundamental attribution error. We explain other people's mistakes by who they are ("careless") and our own by the situation ("the pipeline let me").
Allspaw calls the alternative the second story: the first story is "human error caused the outage"; the second asks what about the system made that error easy, likely or invisible. Google's book says it shortest: "You can't 'fix' people, but you can fix systems and processes."
Blameless is not consequence-free. Accountability moves forward in time: to the action items, which have owners and dates and are tracked until they are done.
Root cause, and why the singular is wrong
A postmortem template usually has a "root cause" heading, and people fill it with one line. Richard Cook's point 7 is blunt: "Post-accident attribution to a 'root cause' is fundamentally wrong." Complex systems fail when several things line up - the release had a slow call, the call had no timeout budget, the pipeline had no canary, the alert fired but the rollback runbook was stale. Remove any one and the incident is smaller or does not happen.
So write contributing factors, plural, and keep asking until you reach ones you can change. Two tools, used with care:
- Five whys (Atlassian's handbook uses it): ask "why?" until you reach something structural. Its weakness: it follows one chain, and stops at the first answer that satisfies the person asking. Ask it more than once, down different branches.
- Proximate vs root, in Atlassian's words: proximate causes "directly led to this incident"; root causes sit "at the optimal place in the chain of events where making a change will prevent this entire class of incident". The second definition is useful precisely because it is about where to act, not who is guilty.
weak: Root cause: Mo deployed 2.9.1 without testing the tax service.
better: Contributing factors:
- 2.9.1 added a synchronous call to the tax service with a 2 s timeout
- the release went to 100% at once: checkout has no canary stage
- the tax service's staging copy answers in 20 ms, production in 1.8-2.5 s
- the rollback runbook referenced a command removed in July
Running the review meeting
The document is drafted before the meeting, by its owner, from the timeline. The meeting is for understanding and for agreeing the actions - usually an hour, within a week of the incident. Someone facilitates who was not the IC (the IC is a participant with a lot to say). A workable agenda:
5 min purpose, ground rules: we are here to understand the system; nobody is on trial
20 min walk the timeline; at each decision: what did you know, what did you expect?
15 min contributing factors, what went well, where we got lucky
15 min action items: owner, date, type, ticket - agreed in the room
5 min who sends the summary, when the actions are reviewed again
The facilitator's main tool is the question. "Why did you deploy on a Friday?" puts a person on the defensive. "What made deploying then look like the right call?" gets the second story: the pipeline was green, the change was small, the team deploys daily. Ask "how" and "what" more than "why you". Refer to people by role ("the on-call", "the release owner") in the document - Atlassian's practice too.
When someone asks "who broke it?" - and in a real meeting someone will, often a manager under pressure - answer the need behind the question: they want to know it will not happen again. "It is not about who; the question is how one config change could reach every customer without a canary. That is what we are fixing, and the owner of that fix is ..." Mission 37.24 puts you in exactly that conversation.
Action items that ship
Most postmortems fail here: twenty items like "improve monitoring" that nobody owns and nobody closes. Atlassian splits follow-ups into priority actions, which carry a deadline (an internal target of 4 or 8 weeks, depending on the service), and the rest. Google's Lueder and Beyer ("Postmortem Action Items: Plan the Work and Work the Plan") want each item actionable (it starts with a verb and says what to do), specific and bounded (it is clear when it is done). The common management acronym says the same thing: SMART - specific, measurable, achievable, relevant, time-bound.
The lab's checker wants five things on every item, plus no non-actions:
owner @team or @person - one name, not "the team"
due a date, YYYY-MM-DD
done when something you can check: a number, a test, an alert that fires
type prevent (it cannot happen again), detect (we see it sooner),
mitigate (it hurts less / is shorter), process (how we work)
ticket the tracking id, so it is followed up like any other work
$ incident check action "Improve monitoring of checkout"
action items (simulator - owner, due date, done-when, type, ticket; no non-actions):
FIX Improve monitoring of checkout
- no owner (@team or @person)
- no due date (YYYY-MM-DD)
- no way to tell it is done (a number, a test, "done when ...")
- no type (prevent / detect / mitigate / process)
- no ticket id
- a non-action ("be more careful", "improve monitoring")
Not actionable yet.
$ incident check action "Add a canary stage to the checkout pipeline that blocks the release when the canary's error rate is over 1% for 5 min - owner: @dana - due 2026-10-20 - type: prevent - CHK-2291"
action items (simulator - owner, due date, done-when, type, ticket; no non-actions):
ok Add a canary stage to the checkout pipeline that blocks the release when the canary's error rate is over 1% for 5 min - owner: @dana - due 2026-10-20 - type: prevent - CHK-2291
All actionable.
What you can now do
- tell mitigated from resolved, and close an incident with the five-point checklist
- name Google's postmortem triggers
- explain just culture, hindsight bias, the fundamental attribution error and the second story
- write contributing factors instead of one root cause, and use five whys with care
- facilitate a review with how/what questions and a timed agenda
- write action items with an owner, a date, a done-when, a type and a ticket