OnCallReady

Lesson 29.15 · Observability III: Dashboards, Logs, Traces & Incidents · 10 min read

Running an incident: severity, command, comms, timeline

In plain words

Imagine a fire at school. If everyone runs around trying to put it out, nobody calls the fire brigade, nobody counts the children, and parents hear rumours. So there is a plan: one teacher is in charge and does not hold a hose; she decides. Some teachers fight the fire, one phones parents every half hour, one writes down what happened and when. First everyone gets outside safely; working out what caused the fire comes later.

That is incident response. The Incident Commander decides severity and roles, responders mitigate, a comms lead updates stakeholders on a cadence, a scribe keeps a UTC timeline. You declare early ("SEV2: checkout failing for ~35%"), mitigate before diagnosing, end when monitoring confirms recovery, then write a blameless postmortem with owned, dated action items.

Why incidents need a process, not just tools

At 20:01 the pager fires. Three people start debugging the same log, nobody tells support, nobody decides to roll back, and a manager asks "any update?" every four minutes. The tools from this chapter find the problem; how long users suffer depends on how the incident is run.

What you need to know already: declaring, severity, the roles (IC, ops, comms, scribe), mitigating first and the update cadence (0.29); blameless postmortems and timelines (0.30, 0.31); error budgets (0.16, 0.17); the burn-rate page OrdersErrorBudgetBurn (28.14); this chapter's dashboard, LogQL and traces (29.1, 29.6, 29.11).

Chapter 0 taught the why and the words. This lesson is the mechanics - the part you practise in the drills and the capstone.

Severity levels

Severity says how bad an incident is, and decides who is woken up, who is told, and how fast (0.29). Every company has its own table; this is 0.29's scale made precise, and the one the drills use (written SEV1, SEV2... here):

SEV1  critical   core user journey down or data loss/security breach, most users
                 affected, no workaround. All hands, IC + comms lead, exec and
                 status-page updates every 30 min.
SEV2  major      significant degradation or a core journey down for a subset of
                 users (one region, one tenant, >5% errors), or a workaround
                 exists but is painful. IC + responders, stakeholder updates hourly.
SEV3  minor      limited impact: a non-core feature broken, elevated errors below
                 SLO-threatening levels, internal tooling down. Handled in hours,
                 by the owning team.
SEV4  low        cosmetic or no user impact yet (a single replica down, a
                 warning trend). A ticket.

A core user journey is something users came to do (log in, pay, check out); a tenant is one customer organisation on a shared system; a replica is one of several identical copies of a service.

Two rules that matter more than the table: declare early (you can downgrade a SEV1 in five minutes; you cannot undo an hour of an undeclared outage), and severity is about user impact, not about how hard the fix is.

Roles

0.29 named them. What each one does, minute to minute:

On a two-person team at night you are all of them. Say which hat you are wearing out loud ("I'm IC, you dig; I'll post the update").

The first ten minutes

  1. Acknowledge the page (so it stops escalating to the next person) and open the incident channel and document.
  2. Assess impact from the SLO dashboard (29.1): which journey, how many users, since when. That sets the severity.
  3. Declare with a one-line summary: "SEV2: checkout failing for ~35% of requests since 20:01, investigating."
  4. Mitigate before you diagnose. Roll back the deploy, stop the load, fail over, shed traffic (0.29). The root cause can wait until users are OK.
  5. Update on a cadence, even if nothing changed ("still investigating, next update 20:30"). Silence is what makes stakeholders escalate.

The timeline

Written during the incident, in UTC, facts only:

20:01  OrdersErrorBudgetBurn fired (1h/5m burn > 14.4x)
20:03  IC: learner. SEV2 declared: checkout 5xx ~35%
20:04  dashboard: errors only on POST /api/checkout, instance localhost:8080
20:06  logs: HikariPool-1 connection timeouts, pool 20/20 active, 40 waiting
20:08  trace 4db9...: 30 s in HikariPool.getConnection, no DB span
20:10  found an unannounced load test (hey, 60 connections) from the QA host
20:11  load test stopped; 5m error ratio back under 0.1% by 20:16
20:40  alert resolved (6h/30m window cleared)

Each line is one time and one fact: what fired, what was declared, what each instrument showed (dashboard, logs, trace - in that order), what was done, and when it was over. 14.4x and the 6h/30m window are the burn-rate alert's thresholds from 28.14; hey is the load generator from 21.17.

The timeline is what the postmortem is built from. Reconstructing it afterwards from chat scrollback is how postmortems get wrong.

Handoffs and ending

A handoff (a shift change, a new IC) is explicit: current impact, current hypothesis, what is in flight, who is doing what, next update time - said out loud and written in the channel. An incident ends when impact is over and monitoring confirms it, not when the fix is merged. Then the IC assigns the postmortem owner and a date.

The postmortem document

0.30 and 0.31 have the why and the blameless language; this is the template the capstone asks for:

# Postmortem: <title>
Status, date, authors, severity

## Summary          two sentences: what happened, impact, how it ended
## Impact           who, how many, how long, SLO budget consumed
## Timeline         from the incident timeline, UTC
## Root cause       the technical chain, without names
## Contributing factors   what made it worse or slower to find
## What went well / what went badly
## Action items     owner, due date, type (prevent / detect / mitigate), ticket

Budget consumed in the Impact section is error-budget arithmetic (0.17): the failed share of requests times the minutes, divided by the month's budget - the next drills practise it. Action items are the only part that changes the future. Each one is specific, owned and dated: not "improve monitoring" but "add a promtool test for CheckoutErrors that reproduces a 45-minute outage - learner

"Walk me through an incident you handled"

The interview version is the postmortem, told in two minutes: the page and the impact, what you checked first and why, the moment you understood it, how you mitigated, the root cause, and the one action item you are proudest of. Have two ready. The capstone of this chapter is built to be one of them.

What you can now do

Why it helps

Tools find the problem; process decides how long users suffer. Most long outages are not hard bugs but chaos: nobody declared, three people debug the same thing, nobody rolls back, stakeholders escalate because they heard nothing. Knowing the mechanics lets you take the IC role calmly on your first real SEV2, even as the most junior person in the channel.

It also matters for your career. "Walk me through an incident you handled" is asked in nearly every SRE interview, and the expected answer has this structure: impact, first checks, the moment you understood it, mitigation, root cause, the action item. With no on-call at your current job, the capstone and the drills in this chapter give you a genuine, well-documented incident to tell. And writing action items like "add a promtool test that reproduces a 45-minute outage, owner, date, ticket" is what reviewers look for.

FAQ

How do I choose the severity?

By user impact, not by how hard the fix looks. Core journey down for most users or data loss is SEV1; significant degradation or a subset of users affected is SEV2; limited impact on non-core features is SEV3; no user impact yet is SEV4. Use your company's table. When in doubt, declare higher: downgrading a SEV1 after five minutes is cheap, while an hour of an undeclared outage cannot be undone.

Why mitigate before finding the root cause?

Because users are hurting now, and mitigation is usually faster and safer than a diagnosis: roll back the last deploy, fail over, stop the offending load, shed traffic, scale out. Once impact is over, you can investigate calmly with the evidence preserved. Keep notes and capture state before mitigating when it is cheap, since a rollback can remove evidence, but do not delay recovery for it.

How often should I send updates?

On a fixed cadence set by severity, such as every 30 minutes for SEV1 and hourly for SEV2, and always say when the next update will come. Send the update even if nothing changed: "still investigating, next update 20:30". Silence is what makes stakeholders escalate and interrupt responders. The communications lead handles this so responders can focus.

When is an incident over?

When user impact has ended and monitoring confirms it, for example the SLO burn alert has resolved and the error ratio is back to normal, not when a fix is merged or deployed. Then the IC announces the end, records the final timeline entries, assigns a postmortem owner and a date, and makes sure any temporary mitigations, silences or disabled automations are tracked for cleanup.

What makes a good postmortem action item?

It is specific, owned, dated and tracked: "add a promtool test for CheckoutErrors that reproduces a 45-minute outage, Learner, 2 October, OPS-2291", not "improve monitoring". Classify it as prevent, detect or mitigate, so you see whether you only fixed this exact cause. Keep the list short enough to actually finish, and review completion in a regular meeting.

In an interview Mid

What are the roles in incident response, and why separate them?

Why separate: without it, three people debug the same log, nobody tells support, nobody decides to roll back, and a manager asks for updates every four minutes. Separation means someone always owns the decisions, the users' impact and the record. On a two-person night shift you wear several hats - say which one out loud.

The mechanics around it: declare early by user impact (SEV1, SEV2...), mitigate before you diagnose, update even when nothing changed, explicit handoffs, and end it when monitoring confirms impact is over - then a blameless postmortem with owned, dated action items.

Also asked: Walk me through an incident you handled. · What do you do in the first ten minutes after being paged? · What makes a good postmortem action item?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.