Why incidents need a process, not just tools
At 20:01 the pager fires. Three people start debugging the same log, nobody tells support, nobody decides to roll back, and a manager asks "any update?" every four minutes. The tools from this chapter find the problem; how long users suffer depends on how the incident is run.
What you need to know already: declaring, severity, the roles (IC, ops, comms, scribe), mitigating first and the update cadence (0.29); blameless postmortems and timelines (0.30, 0.31); error budgets (0.16, 0.17); the burn-rate page OrdersErrorBudgetBurn (28.14); this chapter's dashboard, LogQL and traces (29.1, 29.6, 29.11).
Chapter 0 taught the why and the words. This lesson is the mechanics - the part you practise in the drills and the capstone.
Severity levels
Severity says how bad an incident is, and decides who is woken up, who is told, and how fast (0.29). Every company has its own table; this is 0.29's scale made precise, and the one the drills use (written SEV1, SEV2... here):
SEV1 critical core user journey down or data loss/security breach, most users
affected, no workaround. All hands, IC + comms lead, exec and
status-page updates every 30 min.
SEV2 major significant degradation or a core journey down for a subset of
users (one region, one tenant, >5% errors), or a workaround
exists but is painful. IC + responders, stakeholder updates hourly.
SEV3 minor limited impact: a non-core feature broken, elevated errors below
SLO-threatening levels, internal tooling down. Handled in hours,
by the owning team.
SEV4 low cosmetic or no user impact yet (a single replica down, a
warning trend). A ticket.
A core user journey is something users came to do (log in, pay, check out); a tenant is one customer organisation on a shared system; a replica is one of several identical copies of a service.
Two rules that matter more than the table: declare early (you can downgrade a SEV1 in five minutes; you cannot undo an hour of an undeclared outage), and severity is about user impact, not about how hard the fix is.
Roles
0.29 named them. What each one does, minute to minute:
- Incident Commander (IC) - owns the incident, not the fix. Sets severity, assigns roles, keeps the timeline of decisions, decides when to escalate, when to roll back, when it is over. Does not debug: the moment the IC starts reading stack traces, nobody is steering. The IC's questions: what is the impact, what do we know, what are we trying, who is doing it, when do we next check in.
- Operations / subject-matter responders - investigate and mitigate, and report back to the IC before doing anything risky.
- Communications lead - status page, stakeholders, customer support. Keeps them away from the responders.
- Scribe - the timeline with timestamps: what was observed, decided, done. On a small incident the IC does this.
On a two-person team at night you are all of them. Say which hat you are wearing out loud ("I'm IC, you dig; I'll post the update").
The first ten minutes
- Acknowledge the page (so it stops escalating to the next person) and open the incident channel and document.
- Assess impact from the SLO dashboard (29.1): which journey, how many users, since when. That sets the severity.
- Declare with a one-line summary: "SEV2: checkout failing for ~35% of requests since 20:01, investigating."
- Mitigate before you diagnose. Roll back the deploy, stop the load, fail over, shed traffic (0.29). The root cause can wait until users are OK.
- Update on a cadence, even if nothing changed ("still investigating, next update 20:30"). Silence is what makes stakeholders escalate.
The timeline
Written during the incident, in UTC, facts only:
20:01 OrdersErrorBudgetBurn fired (1h/5m burn > 14.4x)
20:03 IC: learner. SEV2 declared: checkout 5xx ~35%
20:04 dashboard: errors only on POST /api/checkout, instance localhost:8080
20:06 logs: HikariPool-1 connection timeouts, pool 20/20 active, 40 waiting
20:08 trace 4db9...: 30 s in HikariPool.getConnection, no DB span
20:10 found an unannounced load test (hey, 60 connections) from the QA host
20:11 load test stopped; 5m error ratio back under 0.1% by 20:16
20:40 alert resolved (6h/30m window cleared)
Each line is one time and one fact: what fired, what was declared, what each instrument showed (dashboard, logs, trace - in that order), what was done, and when it was over. 14.4x and the 6h/30m window are the burn-rate alert's thresholds from 28.14; hey is the load generator from 21.17.
The timeline is what the postmortem is built from. Reconstructing it afterwards from chat scrollback is how postmortems get wrong.
Handoffs and ending
A handoff (a shift change, a new IC) is explicit: current impact, current hypothesis, what is in flight, who is doing what, next update time - said out loud and written in the channel. An incident ends when impact is over and monitoring confirms it, not when the fix is merged. Then the IC assigns the postmortem owner and a date.
The postmortem document
0.30 and 0.31 have the why and the blameless language; this is the template the capstone asks for:
# Postmortem: <title>
Status, date, authors, severity
## Summary two sentences: what happened, impact, how it ended
## Impact who, how many, how long, SLO budget consumed
## Timeline from the incident timeline, UTC
## Root cause the technical chain, without names
## Contributing factors what made it worse or slower to find
## What went well / what went badly
## Action items owner, due date, type (prevent / detect / mitigate), ticket
Budget consumed in the Impact section is error-budget arithmetic (0.17): the failed share of requests times the minutes, divided by the month's budget - the next drills practise it. Action items are the only part that changes the future. Each one is specific, owned and dated: not "improve monitoring" but "add a promtool test for CheckoutErrors that reproduces a 45-minute outage - learner
- 2 Oct - OPS-2291".
"Walk me through an incident you handled"
The interview version is the postmortem, told in two minutes: the page and the impact, what you checked first and why, the moment you understood it, how you mitigated, the root cause, and the one action item you are proudest of. Have two ready. The capstone of this chapter is built to be one of them.
What you can now do
- Pick a severity from user impact, and declare it in one line.
- Run the first ten minutes: acknowledge, assess, declare, mitigate, update - and keep a UTC timeline while you do.
- Turn that timeline into a postmortem with owned, dated action items.