Why this lesson
The page fires. Three engineers start poking at the problem, each in a private chat, each trying a different fix. Nobody tells customer support what is going on. Forty minutes later two of the fixes have undone each other and nobody can say when it started. Most incidents are not made long by a hard bug; they are made long by nobody being in charge. This lesson is how a team runs one.
What you need to know already: 0.21 Alerting philosophy (pages, on-call, runbooks), 0.16 Error budgets.
Declare early
An incident is not a state of the system, it is a decision: from now on this gets coordinated, not just debugged. The most common failure is not a wrong fix, it is declaring too late - three engineers debugging the same thing in private messages for 40 minutes while support has no idea what to tell customers. Declaring is cheap and reversible (downgrade it, or close it as a false alarm). Not declaring is how incidents drift.
Declare when any of these is true: users are visibly affected; you need help from another team; you cannot see the end within an hour; it is getting worse; or you are simply not sure. Severity (SEV) says how bad it is; a typical scale (every company has its own words):
SEV-1 widespread outage, data loss or security breach; all hands, executives informed
SEV-2 a major feature degraded for many users (checkout failing for a third)
SEV-3 minor or partial impact, a workaround exists, or a single customer
SEV-4 no user impact yet, but needs coordinated work (a near miss: it almost went wrong)
Severity can change during the incident, in both directions. It decides who gets pulled in and how often you communicate, so set it early and adjust.
Roles
From Google's incident management process (itself borrowed from how firefighters organise, the Incident Command System):
- Incident commander (IC) - owns the incident. Keeps the big picture, assigns roles, decides (roll back or not, call in more people or not), makes sure someone owns each open question. The IC does not debug: the moment they open a terminal, nobody is watching the whole.
- Ops lead - the hands on keyboard. The only one changing production (the live system real users are on), so changes do not collide. Says what they are about to do before doing it.
- Comms lead - the status page (the public page that tells customers what is broken), stakeholders, support, executives, on a fixed schedule. Shields the ops lead from "any update?" messages.
- Scribe (or the chat channel itself) - the timeline: every decision, action and observation with a timestamp. This becomes the postmortem.
In a small team one person holds several roles; the IC and ops roles are the two worth splitting as soon as a second person joins. In #4471 the on-call declared SEV-2 and handed IC to a colleague so she could stay on the keyboard - the right move.
Mitigate first
Mitigating means reducing the harm to users - not yet fixing the underlying bug. The goal of an incident is to stop the user impact, not to understand it. Understanding comes afterwards, with time and without pressure. The mitigation menu, roughly from cheapest to most drastic:
roll back the last change the most common fix: most incidents follow a change
turn off a feature flag if the change was behind one (a switch in the
settings that turns a feature on or off)
drain or fail over move traffic away from the broken copy, data centre
or region to healthy ones
scale out add more copies, if it is capacity (and truly is)
shed load refuse some requests, switch off expensive extras
restart buys time; destroys evidence - capture it first
"Is it the release? Can we roll back?" (the IC in #4471) is the first question for a reason. If a change went out in the hour before the impact started, rolling it back is the default, and needs no proof that it is the cause.
Communication
Status updates go out on a fixed cadence (every 30 minutes for a SEV-2), even when there is nothing new - "still investigating, next update at 19:30" is an update. Each one has the same shape:
[SEV-2] checkout: payments failing for about a third of customers
Impact: since 18:32 UTC, ~30% of checkout requests fail with an error
Status: cause identified (release 2.8.0); rollback in progress
Next: update at 19:15 UTC, or sooner if resolved
(UTC is the world reference time zone; incident times are always in UTC so people in different countries agree.) Customers get the impact and the next update time, not the internals. Internal stakeholders get the same, plus what help is needed.
Handoffs (a shift ends, the IC changes) are explicit: "I am handing IC to X. Current state: ..., open actions: ..., next update due at ...". X confirms: "I am IC." Not knowing who is in charge is its own incident.
The clock
Four timestamps summarise the response. From #4471's own data - the first and last failed request from the access log, and the page and its acknowledgement from the pager export:
$ cd ~/oncall-lab/labs/0-sre/incident-4471
$ s() { date -u -d "$1" +%s; }
$ first=$(jq -r 'select(.status >= 500) | .time' /var/log/nginx/checkout.access.log | head -1)
$ last=$(jq -r 'select(.status >= 500) | .time' /var/log/nginx/checkout.access.log | tail -1)
$ page=$(awk '/#4471/ && / TRIGGERED / {print $2}' pager.txt)
$ ack=$(awk '/#4471/ && / ACKNOWLEDGED / {print $2}' pager.txt)
$ echo $first $page $ack $last
2026-09-22T18:32:02+00:00 2026-09-22T18:36:00Z 2026-09-22T18:38:10Z 2026-09-22T19:08:26+00:00
$ echo "detect $(( ($(s $page) - $(s $first)) / 60 )) min, ack $(( $(s $ack) - $(s $page) )) s, mitigate $(( ($(s $last) - $(s $first)) / 60 )) min"
detect 3 min, ack 130 s, mitigate 36 min
The pieces:
first=$(...)stores a command's output in a variable (a named value);$firstreads it back.echoprints its arguments.awk '/#4471/ && / TRIGGERED / {print $2}'prints column 2 (the time) of the pager lines that contain both "#4471" and " TRIGGERED ".s()is a function (0.2):date -u -d "$1" +%sturns a timestamp into seconds since 1 January 1970 (-u= UTC,-d= "this date, not now",+%s= print as seconds). Both the+00:00and theZendings mean UTC.$(( ))does whole-number arithmetic, so 3 min 58 s prints as 3.
The definitions (you will also see them as MTTD, MTTA, MTTR - "mean time to detect / acknowledge / recover or resolve" - when averaged across incidents):
time to detect impact start -> the first alert or report
time to acknowledge alert -> a human says "mine"
time to mitigate impact start -> users stop being hurt
time to resolve impact start -> the underlying problem is fixed (can be days)
For #4471: detection was good (the burn-rate alert, 4 minutes), acknowledgement fine, mitigation slow - about 13 minutes of the 36 were spent discovering that the runbook's deploy rollback command no longer existed. That breakdown points straight at the fix: the time went into recovery, not detection.
A warning about averages: "mean time to X" across incidents is a weak metric. Incident durations are long-tailed (a few huge ones dominate), so the mean jumps around with no change in how well the team works - the same trap as average latency in 0.2. Use the breakdown per incident to find where the time goes; do not set targets on the average.
When it is over
Close when the impact is gone and it is stable (watch for 15-30 minutes after the fix), not when the fix is deployed. Before closing: post the all-clear, name the postmortem owner and the review date, and file the obvious follow-ups while everyone remembers. Google's triggers for a mandatory postmortem include: user-visible impact over a threshold, any data loss, on-call intervention such as a rollback, resolution time over a threshold, and a monitoring failure (a human found it before an alert did). Checkout's policy adds one: more than 20% of the budget in one incident.
In short
declare early, cheaply; set severity, adjust later
roles IC decides and does not debug; ops lead changes production; comms talks
mitigate roll back first, understand later
comms fixed cadence, impact + next update time
clock detect / acknowledge / mitigate / resolve - per incident, not averaged
close stable for a while, all-clear posted, postmortem owner named
What you can now do:
- declare an incident, pick its severity and name the roles
- choose a mitigation before looking for the root of the problem
- compute time to detect, acknowledge and mitigate from logs and a pager export