OnCallReady

Lesson 0.30 · SRE Fundamentals · 9 min read

Blameless postmortems, toil, and running an incident

In plain words

When a football team loses a match, a good coach doesn't say "it was the goalkeeper's fault" and bench him. The coach asks: why was the goalkeeper alone? Why did the defence leave that gap? Why didn't anyone notice the other team's trick earlier? The players then tell the truth about what they saw, because nobody gets punished for honesty, and the team changes its training so it won't happen again.

A blameless postmortem does that for incidents. It assumes everyone acted reasonably given what they knew, uses roles instead of names in the story, looks for several contributing factors instead of one root cause, and ends with action items that are specific, owned and ticketed. Toil, manual repetitive work that grows with the service, is capped at about 50% so there is time to make those fixes.

Why this lesson

After an outage, a manager asks "who did this?". The engineer who pressed the button is named, told to be more careful, and the meeting ends. Three months later a different engineer presses a different button and the same outage happens again - because nothing about the system changed. A postmortem (a written review of an incident: what happened, why, and what will change) is how teams learn from failure instead of repeating it. This lesson covers how to write one that works, and the other half of the SRE job: getting rid of toil.

What you need to know already: 0.29 Running an incident (roles, mitigation, the clock), 0.16 Error budgets.

Blameless, concretely

A blameless postmortem assumes everyone involved acted reasonably given what they knew at the time, and asks how the system made the failure possible - not who failed. Concretely:

Why blame destroys the information

The facts you need live in people's heads: what they saw, what they thought it meant, what they tried and gave up on. If telling the truth gets you blamed:

Blameless is not "nicer". It is the only way to get an accurate account, and the only way the fix lands on the system instead of on a person.

A useful postmortem

Loosely the Google template (written in markdown, the plain-text format where # starts a heading and - starts a bullet):

# Postmortem: <title>            status, date, authors
## Summary                        three lines: what happened, how long, how bad
## Impact                         users, requests, minutes, error budget, money
## Timeline                       UTC, one event per line, from the trigger to all-clear
## Contributing factors           trigger + the conditions that let it hurt
## Detection                      how we found out, how long it took
## Resolution                     what fixed it
## Action items                   action - type - owner - priority - ticket
## Lessons learned                what went well / what went wrong / where we got lucky

Contributing factors, not a root cause. Complex systems fail when several things line up: a change (the trigger), plus the missing safety net, plus the gap in detection, plus the slow recovery. "Root cause: bad config" stops the analysis at the first satisfying answer and leaves three holes open. Ask what made each factor possible.

Action items: real vs theatre

Action items are the follow-up tasks a postmortem creates. Real ones are:

Theatre: "be more careful", "improve monitoring", "raise awareness", "retrain the team", "investigate X" with no output. A good mix covers prevent (stop it happening again), mitigate (smaller blast radius - how much gets hurt - and faster recovery) and detect (find it sooner).

Toil

From the SRE book (chapter 5), toil is work that is:

manual         a human runs it
repetitive     the same thing again and again
automatable    a machine could do it
tactical       interrupt-driven, reactive
no enduring value   the service is no better afterwards
O(n)           grows in step with the service: 2x customers = 2x the work

Not toil: overhead (meetings, planning, 1:1s, training, HR), and engineering (automation, design, writing the SLO, fixing the bug that caused the restarts, writing the postmortem). Unpleasant is not the same as toil: a painful one-off migration has lasting value.

Google caps SRE toil at 50% of time. Reducing it is the job: toil scales with the service and engineering does not, so a team that does not automate ends up needing a new hire for every growth step, with no time left to make anything better. The measure of SRE work is that the service needs fewer humans per unit of growth.

Running an incident (the practice half)

A recap of 0.29, because the postmortem is written from it:

What you can now do:

Why it helps

Postmortems are how SRE teams actually get better, and you'll write, review and present them. Situations: a draft says "Radu skipped the canary"; you rewrite it as "the deploy ran with --skip-canary; the canary stage had been timing out for a week", which leads to a real fix. A postmortem's action items say "be more careful" and "improve monitoring"; you replace them with owned, ticketed changes. A team spends 80% of its time restarting workers; you use the toil definition and the 50% cap to argue for engineering time. Interviewers routinely ask "what does blameless mean?" and "tell me about a postmortem you wrote".

FAQ

Doesn't blameless mean nobody is responsible?

No. Blameless means a person is never treated as the cause, because if a human could make the mistake, the system allowed it: a missing guardrail, a stale runbook, a misleading dashboard. Responsibility moves to fixing: action items have named owners and priorities. And nobody is punished for what a postmortem reveals, which is what keeps people honest.

Why contributing factors instead of one root cause?

Complex systems fail when several defences fail at once: the trigger (a change that shrank the pool), plus a missing safety net (the canary was skipped), plus a detection gap (the health check didn't touch the database), plus slow recovery (a stale runbook). "Root cause: bad config" stops at the first satisfying answer and leaves three holes open. Ask what made each factor possible.

What makes an action item real rather than theatre?

It's specific ("add a startup check that fails if maximumPoolSize is below 20"), owned by a named person, prioritised and dated, tracked in the same tracker as feature work, verifiable (it has a done-when) and about the system. "Be more careful", "improve monitoring" and "retrain the team" are theatre. A good set covers prevent, detect and mitigate.

What exactly counts as toil?

Work that is manual, repetitive, automatable, tactical (interrupt-driven), has no enduring value, and grows linearly with the service. Restarting the same worker every night is toil. Meetings, planning and training are overhead, not toil. Writing an SLO, a postmortem, automation or a painful one-off migration is engineering: unpleasant is not the same as toil if the service is better afterwards.

Why cap toil at 50%?

Because toil grows linearly with the service and engineering doesn't: twice the tenants means twice the manual tickets. A team above the cap stops improving anything and needs a new hire for every growth step. The cap protects time for engineering work that reduces toil, and the measure of SRE work is that the service needs fewer humans per unit of growth.

In an interview Junior

What is a blameless postmortem?

A written review after an incident that asks how the system let the failure happen, not who failed. It assumes everyone acted reasonably with what they knew at the time.

Concretely: it names roles ("the on-call"), not people; it describes what people saw and why their action made sense then; it avoids hindsight words like "should have" and "forgot"; and "human error" is never the cause - if a person could make the mistake, a guardrail was missing. Its sections: summary, impact in numbers (including error budget), a UTC timeline, contributing factors, detection, resolution, action items with owner and ticket, and lessons learned.

Why: the facts live in people's heads. If honesty gets people blamed, the accounts get short and defensive, near misses stop being reported, and the real conditions stay in place for the next person.

Also asked: What makes a postmortem action item good, and what makes one useless? · What is toil? Give an example. · Why does a postmortem include "where we got lucky"?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.