OnCallReady

Lesson 31.28 · Incident Command & Communication · 14 min read

Incident metrics and their traps; on-call health

In plain words

Ten people sit in a cafe, each earning an ordinary salary. Then a billionaire walks in. The average income in the room is suddenly enormous, yet nobody else got a raise; when the billionaire leaves, it drops again. The middle person's income, the median, barely moves at all.

Incident durations behave like that room. Most incidents are short and a few are huge, so the mean time to anything jumps around when nothing changed. In the lab's quarter of twelve incidents, one 290-minute vendor outage drags the mean time to mitigate to 40 minutes while the median is 17. So report the median and the long tail by name, break each incident into detect, acknowledge, coordinate and mitigate, measure what users felt, and never turn one of these numbers into a target. And count the people: pages per shift and nights broken.

Why this lesson

Sooner or later someone asks for a number: "what is our MTTR?", "are incidents getting better?", "how is on-call?". The numbers are easy to compute and easy to misuse. Used badly, they reward the wrong behaviour - closing incidents early, not declaring, splitting one outage into three - and they say "better" or "worse" when nothing changed. Used well, they point at where the time goes and at people who are being worn down.

Chapter 0 (0.29) defined the clock: time to detect, acknowledge, mitigate, resolve. This lesson is about what happens when you average those numbers across incidents, what to look at instead, and how to measure the health of on-call itself.

What you need to know already: 0.29 (the four times), 0.2 (why averages lie for latency - the same trap again), 7.11 (jq).

The MTT-family, and the R that means four things

MTTD   mean time to detect        impact starts -> an alert or a person notices
MTTA   mean time to acknowledge   page -> a human says "mine"
MTTM   mean time to mitigate      impact starts -> users stop being hurt
MTTR   mean time to ...           repair? recover? restore? resolve? respond?

The first trap is the letter R. Different companies, tools and reports use MTTR for "repair", "recovery", "restore" or "resolve", and they are different moments: recovery of service can come hours before the repair of the underlying bug. Two teams comparing MTTRs are often comparing different clocks. Write out which one you mean.

The second trap is the name DORA. Google's DevOps Research and Assessment programme publishes the "DORA metrics" of software delivery; one was "time to restore service". In the 2023 State of DevOps report it was renamed failed deployment recovery time and narrowed to failures caused by a deployment (an outage of a data centre no longer counts). That DORA has nothing to do with the EU regulation called DORA in lesson 37.30.

Why the mean misleads

Incident durations are not spread around an average like people's heights. Most incidents are short; a few are enormous. The Verica Open Incident Database (VOID) report of 2021, built from about 1,800 public incident reports, found durations heavily skewed in exactly this way, and its 2022 report found no correlation between how long an incident lasted and how severe it was. Google's Štěpán Davidovič simulated it in Incident Metrics in SRE (2021): with realistic distributions and incident counts, MTTR and its relatives are "poorly suited for decision making or trend analysis" - a real improvement is lost in the noise, and the number moves when nothing changed.

You can see it on one quarter of checkout incidents. This lesson's setup put the export in ~/oncall-lab/labs/ic/metrics/:

$ cd ~/oncall-lab/labs/ic/metrics
$ jq length incidents.json
12
$ jq -r '.[] | [.id, .sev, .detect_min, .ack_min, .mitigate_min] | @tsv' incidents.json
INC-4398	2	3	2	12
INC-4399	3	11	4	18
INC-4401	2	2	1	9
INC-4402	3	6	9	25
INC-4403	2	4	2	14
INC-4404	3	14	3	31
INC-4405	2	3	2	11
INC-4406	3	5	7	16
INC-4407	1	2	1	290
INC-4410	2	9	2	22
INC-4412	3	4	6	13
INC-4433	2	3	2	19

Columns: incident, severity, minutes to detect, to acknowledge (after the page), to mitigate (from the start of impact). The mean and the median time to mitigate:

$ jq '[.[].mitigate_min] | add / length' incidents.json
40
$ jq '[.[].mitigate_min] | sort' -c incidents.json
[9,11,12,13,14,16,18,19,22,25,31,290]
$ jq '[.[].mitigate_min] | sort | (.[5] + .[6]) / 2' incidents.json
17

add / length is the mean: 40 minutes. The median is the middle of the sorted list - with 12 values, the average of the 6th and 7th (.[5] and .[6], counting from 0): 17 minutes. One incident, INC-4407 at 290 minutes (a vendor outage nobody could fix), more than doubles the mean. Next quarter, without it, "MTTR improved 55%" - and the team did nothing differently.

What to look at instead

And never make one of these numbers a target. Goodhart's law (in Marilyn Strathern's phrasing): "when a measure becomes a target, it ceases to be a good measure." Target MTTR and incidents get resolved too early and reopened as new ones. Target "fewer incidents" and people stop declaring - the exact opposite of 37.2.

On-call health

The people are a system too, and they can be overloaded. Google's SRE book sets the limits its teams run on (Being On-Call, chapter 11):

at least 50% of SRE time on engineering work; of the rest, at most 25% on call
an average of no more than 2 incidents per 12-hour shift (one takes ~6 hours to
    handle well, postmortem included)
on-call teams of at least 8 engineers on one site, or 6 per site on two sites
    (so each person is on call often enough to stay sharp, not so often they burn out)
overload signs: more than 5 tickets a day, 2 or more paging events per shift

Signals worth tracking per rota, each week:

Burnout does not announce itself; it shows as slower acknowledgements, shorter postmortems, people swapping out of rotas, and good engineers leaving. Healthy teams give time off after a bad night, rotate the IC during long incidents (37.14), and treat a noisy pager as a bug with an owner.

In an interview

"How would you measure incident response?" is a common Mid-level question. A strong answer: not by MTTR alone - say which R, show the median and the long tail, break each incident into detect / acknowledge / coordinate / mitigate to see where the time goes, measure user impact with SLOs, and watch on-call load (pages per shift, out of hours).

What you can now do

Why it helps

Sooner or later someone asks "what is our MTTR?" or "are incidents getting better?". The numbers are easy to compute and easy to misuse: used badly they reward closing incidents early, not declaring, or splitting one outage into three, and they say "improved 55%" when the team did nothing differently. Used well, they show where the time goes and who is being worn down.

"How would you measure incident response?" is a common Mid-level interview question, and this lesson is a strong answer: say which R you mean, show the median and the tail, use the per-incident breakdown, measure user impact with SLOs, and watch on-call load against Google's limits. You also practise computing it yourself with jq on an export, which is exactly how you would check a vendor dashboard's claim.

Commands in this lesson

cd jq

FAQ

What does the R in MTTR stand for?

It depends on who wrote it: repair, recovery, restore, resolve, sometimes respond. They are different moments: recovery of service can come hours before the repair of the underlying bug. Two teams comparing MTTRs are often comparing different clocks, so write out which one you mean. MTTD (detect), MTTA (acknowledge) and MTTM (mitigate) are clearer, but averaging them has the same problem.

Why does the mean mislead for incident durations?

Durations are heavily skewed: most incidents are short and a few are enormous. The VOID report of 2021, built from about 1,800 public incident reports, found exactly that, and its 2022 report found no correlation between duration and severity. Štěpán Davidovič's simulations in Incident Metrics in SRE (2021) showed MTTR is poorly suited for decision making or trend analysis: real improvements get lost in noise.

What should I report instead of an average?

Distributions: the median and a high percentile such as the 90th, plus the longest incidents by name ("median 17 min; one over an hour, INC-4407, 290 min, a vendor outage"). The per-incident breakdown of detect, acknowledge, coordinate and mitigate shows what to fix. User impact, such as error budget spent, beats duration. And the reviews: John Allspaw calls the counts shallow data that generate very little insight.

Why not set a target on MTTR or on the number of incidents?

Goodhart's law, in Marilyn Strathern's phrasing: when a measure becomes a target, it ceases to be a good measure. Target MTTR and incidents get resolved too early and reopened as new ones. Target fewer incidents and people stop declaring, which is the opposite of declaring early. Use the numbers to ask questions, not to grade teams.

How can I tell whether on-call is unhealthy?

Compare with Google's limits: at least 50% of SRE time on engineering work and at most 25% on call, on average no more than two incidents per 12-hour shift, teams of at least 8 on one site or 6 per site on two. Track pages per shift and how many were actionable, pages out of hours, time on incidents, and who carries it. Burnout shows as slower acks, shorter postmortems and people leaving rotas.

In an interview Mid

How would you measure incident response?

Not by MTTR alone. First I say which R I mean, because repair, recovery, restore and resolve are different clocks. Then:

And I never make one of these a target: Goodhart's law - target MTTR and incidents get closed early; target fewer incidents and people stop declaring.

Also asked: Why can MTTR improve when the team changed nothing? · What are the DORA metrics of software delivery? · How would you tell that an on-call rota is burning people out?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.