OnCallReady

Lesson 0.29 · SRE Fundamentals · 16 min read

Running an incident: severity, roles, comms, and the clock

In plain words

When there's a fire in a school, the fire brigade doesn't send everyone in to fight it. One chief stands outside and decides: who goes where, whether to call more trucks. Firefighters with hoses do the work. Someone talks to worried parents, so the firefighters aren't interrupted. Someone writes down what happened when. And first they put out the fire; working out how it started comes later.

Running an incident uses the same roles: an incident commander who decides and does not debug, an ops lead who is the only one changing production, a comms lead posting updates on a fixed cadence, and a scribe keeping the timeline. You declare early with a severity (SEV-1 to SEV-4), mitigate first (roll back the last change), and track the clock: time to detect, acknowledge, mitigate and resolve.

Why this lesson

The page fires. Three engineers start poking at the problem, each in a private chat, each trying a different fix. Nobody tells customer support what is going on. Forty minutes later two of the fixes have undone each other and nobody can say when it started. Most incidents are not made long by a hard bug; they are made long by nobody being in charge. This lesson is how a team runs one.

What you need to know already: 0.21 Alerting philosophy (pages, on-call, runbooks), 0.16 Error budgets.

Declare early

An incident is not a state of the system, it is a decision: from now on this gets coordinated, not just debugged. The most common failure is not a wrong fix, it is declaring too late - three engineers debugging the same thing in private messages for 40 minutes while support has no idea what to tell customers. Declaring is cheap and reversible (downgrade it, or close it as a false alarm). Not declaring is how incidents drift.

Declare when any of these is true: users are visibly affected; you need help from another team; you cannot see the end within an hour; it is getting worse; or you are simply not sure. Severity (SEV) says how bad it is; a typical scale (every company has its own words):

SEV-1   widespread outage, data loss or security breach; all hands, executives informed
SEV-2   a major feature degraded for many users (checkout failing for a third)
SEV-3   minor or partial impact, a workaround exists, or a single customer
SEV-4   no user impact yet, but needs coordinated work (a near miss: it almost went wrong)

Severity can change during the incident, in both directions. It decides who gets pulled in and how often you communicate, so set it early and adjust.

Roles

From Google's incident management process (itself borrowed from how firefighters organise, the Incident Command System):

In a small team one person holds several roles; the IC and ops roles are the two worth splitting as soon as a second person joins. In #4471 the on-call declared SEV-2 and handed IC to a colleague so she could stay on the keyboard - the right move.

Mitigate first

Mitigating means reducing the harm to users - not yet fixing the underlying bug. The goal of an incident is to stop the user impact, not to understand it. Understanding comes afterwards, with time and without pressure. The mitigation menu, roughly from cheapest to most drastic:

roll back the last change        the most common fix: most incidents follow a change
turn off a feature flag          if the change was behind one (a switch in the
                                 settings that turns a feature on or off)
drain or fail over               move traffic away from the broken copy, data centre
                                 or region to healthy ones
scale out                        add more copies, if it is capacity (and truly is)
shed load                        refuse some requests, switch off expensive extras
restart                          buys time; destroys evidence - capture it first

"Is it the release? Can we roll back?" (the IC in #4471) is the first question for a reason. If a change went out in the hour before the impact started, rolling it back is the default, and needs no proof that it is the cause.

Communication

Status updates go out on a fixed cadence (every 30 minutes for a SEV-2), even when there is nothing new - "still investigating, next update at 19:30" is an update. Each one has the same shape:

[SEV-2] checkout: payments failing for about a third of customers
Impact:   since 18:32 UTC, ~30% of checkout requests fail with an error
Status:   cause identified (release 2.8.0); rollback in progress
Next:     update at 19:15 UTC, or sooner if resolved

(UTC is the world reference time zone; incident times are always in UTC so people in different countries agree.) Customers get the impact and the next update time, not the internals. Internal stakeholders get the same, plus what help is needed.

Handoffs (a shift ends, the IC changes) are explicit: "I am handing IC to X. Current state: ..., open actions: ..., next update due at ...". X confirms: "I am IC." Not knowing who is in charge is its own incident.

The clock

Four timestamps summarise the response. From #4471's own data - the first and last failed request from the access log, and the page and its acknowledgement from the pager export:

$ cd ~/oncall-lab/labs/0-sre/incident-4471
$ s() { date -u -d "$1" +%s; }
$ first=$(jq -r 'select(.status >= 500) | .time' /var/log/nginx/checkout.access.log | head -1)
$ last=$(jq -r 'select(.status >= 500) | .time' /var/log/nginx/checkout.access.log | tail -1)
$ page=$(awk '/#4471/ && / TRIGGERED / {print $2}' pager.txt)
$ ack=$(awk '/#4471/ && / ACKNOWLEDGED / {print $2}' pager.txt)
$ echo $first $page $ack $last
2026-09-22T18:32:02+00:00 2026-09-22T18:36:00Z 2026-09-22T18:38:10Z 2026-09-22T19:08:26+00:00
$ echo "detect $(( ($(s $page) - $(s $first)) / 60 )) min, ack $(( $(s $ack) - $(s $page) )) s, mitigate $(( ($(s $last) - $(s $first)) / 60 )) min"
detect 3 min, ack 130 s, mitigate 36 min

The pieces:

The definitions (you will also see them as MTTD, MTTA, MTTR - "mean time to detect / acknowledge / recover or resolve" - when averaged across incidents):

time to detect       impact start -> the first alert or report
time to acknowledge  alert -> a human says "mine"
time to mitigate     impact start -> users stop being hurt
time to resolve      impact start -> the underlying problem is fixed (can be days)

For #4471: detection was good (the burn-rate alert, 4 minutes), acknowledgement fine, mitigation slow - about 13 minutes of the 36 were spent discovering that the runbook's deploy rollback command no longer existed. That breakdown points straight at the fix: the time went into recovery, not detection.

A warning about averages: "mean time to X" across incidents is a weak metric. Incident durations are long-tailed (a few huge ones dominate), so the mean jumps around with no change in how well the team works - the same trap as average latency in 0.2. Use the breakdown per incident to find where the time goes; do not set targets on the average.

When it is over

Close when the impact is gone and it is stable (watch for 15-30 minutes after the fix), not when the fix is deployed. Before closing: post the all-clear, name the postmortem owner and the review date, and file the obvious follow-ups while everyone remembers. Google's triggers for a mandatory postmortem include: user-visible impact over a threshold, any data loss, on-call intervention such as a rollback, resolution time over a threshold, and a monitoring failure (a human found it before an alert did). Checkout's policy adds one: more than 20% of the budget in one incident.

In short

declare     early, cheaply; set severity, adjust later
roles       IC decides and does not debug; ops lead changes production; comms talks
mitigate    roll back first, understand later
comms       fixed cadence, impact + next update time
clock       detect / acknowledge / mitigate / resolve - per incident, not averaged
close       stable for a while, all-clear posted, postmortem owner named

What you can now do:

Why it helps

You will run incidents, and the difference between a 36-minute outage and a 3-hour one is often coordination, not technical skill. Situations: three engineers debugging checkout in private DMs for 40 minutes while support has nothing to tell customers, because nobody declared. The IC (incident commander) starts typing commands on the servers and nobody is watching the whole picture. A release went out an hour before the errors began, and the team spends 30 minutes proving it's the cause instead of rolling back. Stakeholders keep interrupting the ops lead for updates. Knowing the roles, the mitigation menu and a status update template makes you useful from your first incident, and interviewers often ask "tell me how you handle an incident".

Commands in this lesson

cd echo

FAQ

When should I declare an incident?

When users are visibly affected, you need help from another team, you can't see the end within an hour, it's getting worse, or you're simply not sure. Declaring is cheap and reversible: you can downgrade or close it as a false alarm. The most common failure isn't a wrong fix, it's declaring too late, with people debugging privately while nobody coordinates or communicates.

Why must the incident commander not debug?

The IC's job is the big picture: assigning roles, tracking open questions, deciding whether to roll back or escalate, and making sure communication happens. The moment the IC opens a terminal, their attention narrows to one hypothesis and nobody is watching the whole. In a small team one person holds several roles, but IC and ops lead are the first two to split when a second person joins.

Should I find the root cause before rolling back?

No. The goal of an incident is to stop user impact; understanding comes later, with time and without pressure. If a change went out in the hour before impact began, rolling it back is the default and needs no proof that it's the cause. Capture evidence first if a mitigation destroys it, for example before restarting.

What goes into a status update?

The same shape every time: severity and a one-line summary, the impact in user terms and since when (UTC), the current status, and the time of the next update. Updates go out on a fixed cadence, every 30 minutes for a SEV-2, even when nothing changed: "still investigating, next update at 19:30" is an update. Customers get impact and timing, not internals.

Why not track mean time to recovery across incidents?

Incident durations are long-tailed: a few huge incidents dominate the mean, so it jumps around without any change in how well the team works. Use the per-incident breakdown (detect, acknowledge, mitigate, resolve) to find where the time went, like 13 minutes lost to a stale runbook in #4471, and fix that. Don't set targets on the average.

In an interview Junior

What roles are there in incident response, and why?

Without roles everyone debugs, nobody talks to customers and nobody can say later what happened when. In a small team one person holds several roles; the first split is IC from ops lead, and handovers are said out loud.

Also asked: What does a SEV-1 mean compared with a SEV-3? · Why mitigate first and look for the cause later? · What goes into a status update during an incident?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.