Why this lesson
Sooner or later someone asks for a number: "what is our MTTR?", "are incidents getting better?", "how is on-call?". The numbers are easy to compute and easy to misuse. Used badly, they reward the wrong behaviour - closing incidents early, not declaring, splitting one outage into three - and they say "better" or "worse" when nothing changed. Used well, they point at where the time goes and at people who are being worn down.
Chapter 0 (0.29) defined the clock: time to detect, acknowledge, mitigate, resolve. This lesson is about what happens when you average those numbers across incidents, what to look at instead, and how to measure the health of on-call itself.
What you need to know already: 0.29 (the four times), 0.2 (why averages lie for latency - the same trap again), 7.11 (jq).
The MTT-family, and the R that means four things
MTTD mean time to detect impact starts -> an alert or a person notices
MTTA mean time to acknowledge page -> a human says "mine"
MTTM mean time to mitigate impact starts -> users stop being hurt
MTTR mean time to ... repair? recover? restore? resolve? respond?
The first trap is the letter R. Different companies, tools and reports use MTTR for "repair", "recovery", "restore" or "resolve", and they are different moments: recovery of service can come hours before the repair of the underlying bug. Two teams comparing MTTRs are often comparing different clocks. Write out which one you mean.
The second trap is the name DORA. Google's DevOps Research and Assessment programme publishes the "DORA metrics" of software delivery; one was "time to restore service". In the 2023 State of DevOps report it was renamed failed deployment recovery time and narrowed to failures caused by a deployment (an outage of a data centre no longer counts). That DORA has nothing to do with the EU regulation called DORA in lesson 37.30.
Why the mean misleads
Incident durations are not spread around an average like people's heights. Most incidents are short; a few are enormous. The Verica Open Incident Database (VOID) report of 2021, built from about 1,800 public incident reports, found durations heavily skewed in exactly this way, and its 2022 report found no correlation between how long an incident lasted and how severe it was. Google's Štěpán Davidovič simulated it in Incident Metrics in SRE (2021): with realistic distributions and incident counts, MTTR and its relatives are "poorly suited for decision making or trend analysis" - a real improvement is lost in the noise, and the number moves when nothing changed.
You can see it on one quarter of checkout incidents. This lesson's setup put the export in ~/oncall-lab/labs/ic/metrics/:
$ cd ~/oncall-lab/labs/ic/metrics
$ jq length incidents.json
12
$ jq -r '.[] | [.id, .sev, .detect_min, .ack_min, .mitigate_min] | @tsv' incidents.json
INC-4398 2 3 2 12
INC-4399 3 11 4 18
INC-4401 2 2 1 9
INC-4402 3 6 9 25
INC-4403 2 4 2 14
INC-4404 3 14 3 31
INC-4405 2 3 2 11
INC-4406 3 5 7 16
INC-4407 1 2 1 290
INC-4410 2 9 2 22
INC-4412 3 4 6 13
INC-4433 2 3 2 19
Columns: incident, severity, minutes to detect, to acknowledge (after the page), to mitigate (from the start of impact). The mean and the median time to mitigate:
$ jq '[.[].mitigate_min] | add / length' incidents.json
40
$ jq '[.[].mitigate_min] | sort' -c incidents.json
[9,11,12,13,14,16,18,19,22,25,31,290]
$ jq '[.[].mitigate_min] | sort | (.[5] + .[6]) / 2' incidents.json
17
add / length is the mean: 40 minutes. The median is the middle of the sorted list - with 12 values, the average of the 6th and 7th (.[5] and .[6], counting from 0): 17 minutes. One incident, INC-4407 at 290 minutes (a vendor outage nobody could fix), more than doubles the mean. Next quarter, without it, "MTTR improved 55%" - and the team did nothing differently.
What to look at instead
- Distributions, not means. The median and a high percentile (the 90th), plus the list of the longest incidents by name. "Median 17 min; one over an hour, INC-4407 (290 min, a vendor outage)" says more than "MTTR 40".
- Where the time went, per incident. The 0.29 breakdown - detect, acknowledge, coordinate, mitigate - points at the fix: slow detection wants better alerts, slow acknowledgement wants a better rota, slow coordination wants incident command, slow mitigation wants better rollbacks.
- User impact, not duration. A 40-minute SEV3 and a 4-minute total outage are not comparable. Error budget spent (chapter 0) measures what users felt.
- The story. John Allspaw calls duration, frequency and severity counts "shallow data": valid, but they "tend to generate very little insight". The insight is in the reviews: what surprised people, where coordination broke, what made the fix slow.
And never make one of these numbers a target. Goodhart's law (in Marilyn Strathern's phrasing): "when a measure becomes a target, it ceases to be a good measure." Target MTTR and incidents get resolved too early and reopened as new ones. Target "fewer incidents" and people stop declaring - the exact opposite of 37.2.
On-call health
The people are a system too, and they can be overloaded. Google's SRE book sets the limits its teams run on (Being On-Call, chapter 11):
at least 50% of SRE time on engineering work; of the rest, at most 25% on call
an average of no more than 2 incidents per 12-hour shift (one takes ~6 hours to
handle well, postmortem included)
on-call teams of at least 8 engineers on one site, or 6 per site on two sites
(so each person is on call often enough to stay sharp, not so often they burn out)
overload signs: more than 5 tickets a day, 2 or more paging events per shift
Signals worth tracking per rota, each week:
- pages per shift, and how many were actionable (a page nobody needed to act on is a broken alert - chapter 28);
- pages out of hours: each one is a broken night of sleep;
- time spent on incidents and follow-ups against the 25% line;
- who carries it: one expert who "always gets pulled in" is a single point of failure and a burnout risk, whatever the rota says.
Burnout does not announce itself; it shows as slower acknowledgements, shorter postmortems, people swapping out of rotas, and good engineers leaving. Healthy teams give time off after a bad night, rotate the IC during long incidents (37.14), and treat a noisy pager as a bug with an owner.
In an interview
"How would you measure incident response?" is a common Mid-level question. A strong answer: not by MTTR alone - say which R, show the median and the long tail, break each incident into detect / acknowledge / coordinate / mitigate to see where the time goes, measure user impact with SLOs, and watch on-call load (pages per shift, out of hours).
What you can now do
- say which clock each MTT-metric measures, and why MTTR is ambiguous
- tell the DevOps DORA metrics from the EU regulation
- show with jq why a mean of skewed durations misleads, and report the median and the tail
- choose better signals (breakdown, user impact, the reviews) and avoid targets
- check on-call load against Google's limits