Why this lesson
Ask four people what happened during an outage and you get four timelines, each with the storyteller's own actions looking more sensible than they were. The postmortem then argues about memories instead of fixing the system. This lesson builds the timeline from what the machines recorded, and then turns it into blameless sentences.
What you need to know already: 0.30 Blameless postmortems (sections, action items, contributing factors), 0.29 Running an incident (the clock, UTC).
Evidence first
A postmortem is built from what the systems recorded, then annotated with what people saw and thought. Memory is the least reliable source: people remember the order of events wrongly, and they remember their own actions as more reasonable than the chat log shows (everyone does). Collect first:
access logs when users were hurt: first and last failed request, per minute
error logs what the proxy / service said was wrong
journal deploys, sudo commands (who ran what, when - as evidence, not blame)
pager export when alerts fired, were acked, resolved
chat export what people knew and decided, and when
The journal is the server's central log: every program on it (including the deploy tool, and sudo, which records each command run as root) sends its messages there, and journalctl reads it. journalctl -t deploy shows only messages tagged "deploy" (-t = tag); -o short-iso prints times in ISO form; --no-pager prints everything at once instead of page by page.
One timeline from many sources
Every source writes its timestamps differently. Cut each one down to YYYY-MM-DDTHH:MM:SS (19 characters), tag it with its source, and let sort merge them:
$ cd ~/oncall-lab/labs/0-sre/incident-4471
$ { journalctl -t deploy -t sudo -o short-iso --no-pager | grep -E 'checkout|rollback|undo' | awk '{t=substr($1,1,19); $1=""; $2=""; print t " journal" $0}'; jq -r 'select(.status >= 500) | .time[0:19]' /var/log/nginx/checkout.access.log | sed -n '1p;$p' | sed 's/$/ nginx 5xx first or last/'; awk '/^#4471/ {print substr($2,1,19), "pager ", $3, $4}' pager.txt; } | sort | cut -c1-110
2026-09-22T18:27:55 journal sudo[3391]: radu : TTY=pts/2 ; PWD=/home/radu ; USER=root ; COMMAND=/usr/local/bi
2026-09-22T18:27:56 journal deploy[3394]: checkout: canary stage skipped (--skip-canary)
2026-09-22T18:27:56 journal deploy[3394]: checkout: deploying 2.8.0 (current: 2.7.3), strategy=rolling, repli
2026-09-22T18:28:51 journal deploy[3394]: checkout: replica 1/3 on 2.8.0 Ready
2026-09-22T18:29:46 journal deploy[3394]: checkout: replica 2/3 on 2.8.0 Ready
2026-09-22T18:30:41 journal deploy[3394]: checkout: replica 3/3 on 2.8.0 Ready
2026-09-22T18:30:42 journal deploy[3394]: checkout: rollout complete in 166s
2026-09-22T18:32:02 nginx 5xx first or last
2026-09-22T18:36:00 pager TRIGGERED CheckoutErrorBudgetBurnFast
2026-09-22T18:38:10 pager ACKNOWLEDGED by
2026-09-22T18:41:20 pager NOTE SEV-2
2026-09-22T18:52:12 journal deploy[4023]: deploy: unknown command 'rollback'
2026-09-22T18:52:12 journal sudo[4020]: ioana : TTY=pts/3 ; PWD=/home/ioana ; USER=root ; COMMAND=/usr/local/
2026-09-22T19:05:04 journal deploy[4191]: checkout: undo 2.8.0 -> 2.7.3, strategy=rolling, replicas=3
2026-09-22T19:05:04 journal sudo[4188]: ioana : TTY=pts/3 ; PWD=/home/ioana ; USER=root ; COMMAND=/usr/local/
2026-09-22T19:06:06 journal deploy[4191]: checkout: replica 1/3 on 2.7.3 Ready
2026-09-22T19:07:05 journal deploy[4191]: checkout: replica 2/3 on 2.7.3 Ready
2026-09-22T19:08:23 journal deploy[4191]: checkout: replica 3/3 on 2.7.3 Ready
2026-09-22T19:08:24 journal deploy[4191]: checkout: rollout complete in 200s
2026-09-22T19:08:26 nginx 5xx first or last
2026-09-22T19:14:00 pager RESOLVED auto-resolved
It looks long; it is three small commands glued together:
{ A; B; C; } | sortruns A, B and C one after another and sends all their output, as one stream, intosort.- A: the journal's deploy and sudo messages;
grep -E 'checkout|rollback|undo'keeps lines containing any of the three words (-Elets|mean "or"); awk keeps the first 19 characters of the time and drops the next column. - B: the first and last 5xx from the access log (
sed -n '1p;$p'prints line 1 and the last line), with a label glued on the end (s/$/ .../: replace the end of the line with the text). - C: the pager lines for #4471.
cut -c1-110keeps the first 110 characters of each line, so it fits.
ISO timestamps sort correctly as plain text - that is the whole trick. Each "replica" line is one copy of checkout switching to the new version. Now the story is readable: deploy at 18:27, first 5xx 80 seconds after the rollout completed, page at 18:36, a failed rollback at 18:52, the working one at 19:05, last 5xx at 19:08.
The gaps are the findings: 11 minutes between the SEV-2 and the first rollback attempt, and 13 minutes between the failed rollback and the undo. The chat explains the second gap (the runbook was out of date); a postmortem that only lists events without asking about the gaps misses the point.
Timeline rules
- UTC, always, and say so. Local time in a timeline written by people in two countries is a guaranteed error.
- Impact start is the first failed request, not when someone noticed. The first draft of #4488 (you review it later) started the clock at the support ticket, 19 minutes late.
- One event per line, in order:
- HH:MM what happened. - Include decisions and observations, not only actions: "18:40 every copy of checkout reports healthy; /healthz passes" explains why nobody looked at the health check.
Blameless, sentence by sentence
Blameless writing is a skill you practise on sentences. The pattern: replace the person with a role, replace the verdict with the condition that made the action reasonable.
BLAME BLAMELESS
Radu skipped the canary. The deploy ran with --skip-canary; the canary
stage had been timing out for a week.
Ioana ran the wrong rollback command. The runbook said 'deploy rollback'; deploy v3
renamed it to 'undo' and the runbook was not updated.
Nobody noticed the pool size change. The new configuration file did not carry
maximumPoolSize over; nothing checks pool settings.
Sorin should have looked at the error log. The upstream certificate error was only in the edge
error log, which is not on the api dashboard.
Human error: the migration ran on prod. The test and production systems had the same name
in the tool, and the prompt did not show which one
was active.
The on-call was slow to respond. The page was acknowledged 12 minutes after it fired;
the phone's focus mode muted the notification.
Words that give blame away: forgot, failed to, should have, careless, missed, ignored, finally, obviously, just, simply, human error, fault. "Should have" is hindsight: it judges a decision with information the person did not have at the time. "Human error" is never a root cause; it is where the analysis stopped.
Blameless does not mean nameless everywhere. Names belong in the action items as owners, and people can be named as sources ("the deploy tool's author explained..."). What they may not be is a cause.
Contributing factors, not a root cause
Complex systems fail when several defences fail at once (James Reason's Swiss cheese model: each layer of defence has holes; an incident is when the holes line up). Sort the factors into the layers:
trigger the change or event that started it 2.8.0 dropped the pool size
prevention what should have stopped it reaching users canary skipped; no config check
detection what should have noticed sooner /healthz did not touch the database
response what slowed recovery stale runbook: rollback -> undo
"Five whys" (asking "why?" again at each answer) is useful for walking down one branch ("why did the pool shrink? why did the change miss it? why is there no check?"), but it pretends there is a single chain. Ask it for every layer, and stop at things you can change, not at a person.
Action items that get done
THEATRE REAL
Be more careful with deploys. Make --skip-canary require a second approver
in the deploy tool - owner: radu.c - P0 - OPS-7101
Improve monitoring. Add a ticket alert when pool requests are pending
for 2 minutes - owner: mihai.d - P1 - OPS-7102
Update the runbooks. Generate runbook commands from 'deploy --help' in the
build, so a renamed command fails it - owner: andrei.p - P1
Investigate the health check. /healthz borrows a database connection; done when a copy
with a broken pool reports unhealthy - owner: ioana.m - P1
(P0, P1 = priority levels, P0 most urgent; OPS-7101 is the ticket number.) Every real one is specific, owned by a person, prioritised, has a ticket in the same tracker as feature work, and has a "done when". Aim for a mix: at least one prevent, one detect, one mitigate. Review open postmortem actions in a standing meeting until they are closed - the postmortem is not finished when the document is, it is finished when its P0s have shipped.
Review checklist
Before the review meeting, read the draft and check:
[ ] impact start and end come from data, in UTC
[ ] impact has numbers: requests or users, minutes, % of error budget
[ ] timeline explains the gaps, not just the events
[ ] three or more contributing factors, across the layers
[ ] no person is a cause; no "should have", "forgot", "human error"
[ ] detection: how did we find out, and should an alert have fired?
[ ] every action item: specific, owner, priority, ticket, done-when
[ ] "where we got lucky" is filled in - luck is a hidden risk
In short
evidence logs, journal, pager, chat - memory last
timeline cut to ISO, tag, sort; read the gaps
blameless role instead of person, condition instead of verdict
factors trigger, prevention, detection, response
actions specific, owned, ticketed, done-when; prevent + detect + mitigate
What you can now do:
- merge several logs into one timeline and read the gaps
- rewrite a blaming sentence as a blameless one
- sort contributing factors into trigger, prevention, detection and response