OnCallReady

Lesson 0.31 · SRE Fundamentals · 18 min read

Writing the postmortem: evidence, timeline, blameless language

In plain words

Imagine working out what happened at a birthday party where the cake fell over. You could ask everyone what they remember, but people remember things in the wrong order and remember themselves as a bit more sensible than they were. Better: look at the photos from everyone's phones, sort them by the time they were taken, and then ask people what they were thinking at each moment. The interesting parts are often the gaps: why did ten minutes pass between the wobble and anyone catching it?

Writing a postmortem works the same way: collect evidence from the access log, journal, pager and chat, normalise every timestamp to ISO UTC, tag and sort them into one timeline, and read the gaps. Then write blamelessly, sentence by sentence: replace the person with a role and the verdict with the condition.

Why this lesson

Ask four people what happened during an outage and you get four timelines, each with the storyteller's own actions looking more sensible than they were. The postmortem then argues about memories instead of fixing the system. This lesson builds the timeline from what the machines recorded, and then turns it into blameless sentences.

What you need to know already: 0.30 Blameless postmortems (sections, action items, contributing factors), 0.29 Running an incident (the clock, UTC).

Evidence first

A postmortem is built from what the systems recorded, then annotated with what people saw and thought. Memory is the least reliable source: people remember the order of events wrongly, and they remember their own actions as more reasonable than the chat log shows (everyone does). Collect first:

access logs        when users were hurt: first and last failed request, per minute
error logs         what the proxy / service said was wrong
journal            deploys, sudo commands (who ran what, when - as evidence, not blame)
pager export       when alerts fired, were acked, resolved
chat export        what people knew and decided, and when

The journal is the server's central log: every program on it (including the deploy tool, and sudo, which records each command run as root) sends its messages there, and journalctl reads it. journalctl -t deploy shows only messages tagged "deploy" (-t = tag); -o short-iso prints times in ISO form; --no-pager prints everything at once instead of page by page.

One timeline from many sources

Every source writes its timestamps differently. Cut each one down to YYYY-MM-DDTHH:MM:SS (19 characters), tag it with its source, and let sort merge them:

$ cd ~/oncall-lab/labs/0-sre/incident-4471
$ { journalctl -t deploy -t sudo -o short-iso --no-pager | grep -E 'checkout|rollback|undo' | awk '{t=substr($1,1,19); $1=""; $2=""; print t " journal" $0}'; jq -r 'select(.status >= 500) | .time[0:19]' /var/log/nginx/checkout.access.log | sed -n '1p;$p' | sed 's/$/ nginx  5xx first or last/'; awk '/^#4471/ {print substr($2,1,19), "pager  ", $3, $4}' pager.txt; } | sort | cut -c1-110
2026-09-22T18:27:55 journal  sudo[3391]: radu : TTY=pts/2 ; PWD=/home/radu ; USER=root ; COMMAND=/usr/local/bi
2026-09-22T18:27:56 journal  deploy[3394]: checkout: canary stage skipped (--skip-canary)
2026-09-22T18:27:56 journal  deploy[3394]: checkout: deploying 2.8.0 (current: 2.7.3), strategy=rolling, repli
2026-09-22T18:28:51 journal  deploy[3394]: checkout: replica 1/3 on 2.8.0 Ready
2026-09-22T18:29:46 journal  deploy[3394]: checkout: replica 2/3 on 2.8.0 Ready
2026-09-22T18:30:41 journal  deploy[3394]: checkout: replica 3/3 on 2.8.0 Ready
2026-09-22T18:30:42 journal  deploy[3394]: checkout: rollout complete in 166s
2026-09-22T18:32:02 nginx  5xx first or last
2026-09-22T18:36:00 pager   TRIGGERED CheckoutErrorBudgetBurnFast
2026-09-22T18:38:10 pager   ACKNOWLEDGED by
2026-09-22T18:41:20 pager   NOTE SEV-2
2026-09-22T18:52:12 journal  deploy[4023]: deploy: unknown command 'rollback'
2026-09-22T18:52:12 journal  sudo[4020]: ioana : TTY=pts/3 ; PWD=/home/ioana ; USER=root ; COMMAND=/usr/local/
2026-09-22T19:05:04 journal  deploy[4191]: checkout: undo 2.8.0 -> 2.7.3, strategy=rolling, replicas=3
2026-09-22T19:05:04 journal  sudo[4188]: ioana : TTY=pts/3 ; PWD=/home/ioana ; USER=root ; COMMAND=/usr/local/
2026-09-22T19:06:06 journal  deploy[4191]: checkout: replica 1/3 on 2.7.3 Ready
2026-09-22T19:07:05 journal  deploy[4191]: checkout: replica 2/3 on 2.7.3 Ready
2026-09-22T19:08:23 journal  deploy[4191]: checkout: replica 3/3 on 2.7.3 Ready
2026-09-22T19:08:24 journal  deploy[4191]: checkout: rollout complete in 200s
2026-09-22T19:08:26 nginx  5xx first or last
2026-09-22T19:14:00 pager   RESOLVED auto-resolved

It looks long; it is three small commands glued together:

ISO timestamps sort correctly as plain text - that is the whole trick. Each "replica" line is one copy of checkout switching to the new version. Now the story is readable: deploy at 18:27, first 5xx 80 seconds after the rollout completed, page at 18:36, a failed rollback at 18:52, the working one at 19:05, last 5xx at 19:08.

The gaps are the findings: 11 minutes between the SEV-2 and the first rollback attempt, and 13 minutes between the failed rollback and the undo. The chat explains the second gap (the runbook was out of date); a postmortem that only lists events without asking about the gaps misses the point.

Timeline rules

Blameless, sentence by sentence

Blameless writing is a skill you practise on sentences. The pattern: replace the person with a role, replace the verdict with the condition that made the action reasonable.

BLAME                                          BLAMELESS
Radu skipped the canary.                       The deploy ran with --skip-canary; the canary
                                               stage had been timing out for a week.
Ioana ran the wrong rollback command.          The runbook said 'deploy rollback'; deploy v3
                                               renamed it to 'undo' and the runbook was not updated.
Nobody noticed the pool size change.           The new configuration file did not carry
                                               maximumPoolSize over; nothing checks pool settings.
Sorin should have looked at the error log.     The upstream certificate error was only in the edge
                                               error log, which is not on the api dashboard.
Human error: the migration ran on prod.        The test and production systems had the same name
                                               in the tool, and the prompt did not show which one
                                               was active.
The on-call was slow to respond.               The page was acknowledged 12 minutes after it fired;
                                               the phone's focus mode muted the notification.

Words that give blame away: forgot, failed to, should have, careless, missed, ignored, finally, obviously, just, simply, human error, fault. "Should have" is hindsight: it judges a decision with information the person did not have at the time. "Human error" is never a root cause; it is where the analysis stopped.

Blameless does not mean nameless everywhere. Names belong in the action items as owners, and people can be named as sources ("the deploy tool's author explained..."). What they may not be is a cause.

Contributing factors, not a root cause

Complex systems fail when several defences fail at once (James Reason's Swiss cheese model: each layer of defence has holes; an incident is when the holes line up). Sort the factors into the layers:

trigger      the change or event that started it          2.8.0 dropped the pool size
prevention   what should have stopped it reaching users   canary skipped; no config check
detection    what should have noticed sooner              /healthz did not touch the database
response     what slowed recovery                         stale runbook: rollback -> undo

"Five whys" (asking "why?" again at each answer) is useful for walking down one branch ("why did the pool shrink? why did the change miss it? why is there no check?"), but it pretends there is a single chain. Ask it for every layer, and stop at things you can change, not at a person.

Action items that get done

THEATRE                               REAL
Be more careful with deploys.         Make --skip-canary require a second approver
                                      in the deploy tool - owner: radu.c - P0 - OPS-7101
Improve monitoring.                   Add a ticket alert when pool requests are pending
                                      for 2 minutes - owner: mihai.d - P1 - OPS-7102
Update the runbooks.                  Generate runbook commands from 'deploy --help' in the
                                      build, so a renamed command fails it - owner: andrei.p - P1
Investigate the health check.         /healthz borrows a database connection; done when a copy
                                      with a broken pool reports unhealthy - owner: ioana.m - P1

(P0, P1 = priority levels, P0 most urgent; OPS-7101 is the ticket number.) Every real one is specific, owned by a person, prioritised, has a ticket in the same tracker as feature work, and has a "done when". Aim for a mix: at least one prevent, one detect, one mitigate. Review open postmortem actions in a standing meeting until they are closed - the postmortem is not finished when the document is, it is finished when its P0s have shipped.

Review checklist

Before the review meeting, read the draft and check:

[ ] impact start and end come from data, in UTC
[ ] impact has numbers: requests or users, minutes, % of error budget
[ ] timeline explains the gaps, not just the events
[ ] three or more contributing factors, across the layers
[ ] no person is a cause; no "should have", "forgot", "human error"
[ ] detection: how did we find out, and should an alert have fired?
[ ] every action item: specific, owner, priority, ticket, done-when
[ ] "where we got lucky" is filled in - luck is a hidden risk

In short

evidence     logs, journal, pager, chat - memory last
timeline     cut to ISO, tag, sort; read the gaps
blameless    role instead of person, condition instead of verdict
factors      trigger, prevention, detection, response
actions      specific, owned, ticketed, done-when; prevent + detect + mitigate

What you can now do:

Why it helps

You will build incident timelines from raw logs under time pressure, and this lesson gives you the one-liner technique: substr($1,1,19), tag, sort. Situations: a draft postmortem starts the impact clock at the support ticket, 19 minutes after the first failed request; the access log proves otherwise. A timeline lists events but not the 13-minute gap between the failed rollback and the undo, which is where the real finding (a stale runbook) was. A colleague's draft says "Ioana ran the wrong rollback command"; you rewrite it as "the runbook said 'deploy rollback'; deploy v3 renamed it to 'undo'". Reviewing postmortems well is a visible senior skill.

Commands in this lesson

cd

FAQ

Why build the timeline from logs instead of asking people?

Memory is the least reliable source: people remember the order of events wrongly and their own actions as more reasonable than the chat log shows. Systems record exact times: first and last failed request from access logs, deploys and sudo commands from the journal, alert times from the pager, decisions from chat. Build the skeleton from evidence, then annotate it with what people saw and thought.

When does the impact clock start?

At the first failed request, from the data, not when someone noticed. In #4488 the draft started at the support ticket, 19 minutes late, which also hid a detection problem: no alert fired. The detection gap is itself one of the most important findings, so the start time must be accurate.

Why always UTC?

Timelines are often assembled from systems and people in different time zones, and logs come in several formats (+00:00, Z, local time). Mixing local times in one timeline is a guaranteed error. Normalise everything to ISO UTC; ISO timestamps also sort correctly as plain text, which is what makes the sort-merge trick work.

Does blameless mean I can never name anyone?

Names belong in action items as owners and can appear as sources ("the deploy tool's author explained..."). What a person may never be is a cause. In the narrative, use roles, and describe the condition that made the action reasonable: not "the on-call was slow" but "the page was acknowledged 12 minutes after firing; the phone's focus mode muted it".

Is "five whys" a good technique?

Useful for walking down one branch, like why the pool shrank, why the migration missed it, why there is no check. But it assumes a single chain, while incidents have several contributing factors. Ask it for every layer (trigger, prevention, detection, response) and stop at things you can change, never at a person.

In an interview Junior

How do you build an incident timeline?

From evidence first, memory last.

I collect the sources: the access log (first and last failed request = when users were hurt), the error logs, the journal (deploys, sudo commands), the pager export (fired, acknowledged, resolved) and the chat export (what people knew and decided). Each uses its own time format, so I convert every line to the same ISO form in UTC, tag it with its source, and merge them with sort - ISO timestamps sort correctly as text.

Rules: always UTC; impact starts at the first failed request, not when someone noticed; one event per line; include decisions and observations, not only actions. Then I read the gaps: in #4471 there were 11 minutes between declaring SEV-2 and the first rollback attempt. Explaining the gaps is where the findings are.

Also asked: Rewrite "the engineer forgot to update the runbook" in blameless language. · What is the difference between a trigger and a contributing factor? · Why is "human error" not accepted as a cause?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.