Why this lesson
A burn-rate alert fires at 02:10. The on-call opens the dashboard, sees a flat line over the last five minutes, decides it is a glitch, and snoozes it. Six hours later a quarter of the month's budget is gone. The alert was right; the way it was designed and shown made it unbelievable. This lesson is about tuning an alert so people trust it: how fast it fires, how fast it clears, and how noisy it is.
What you need to know already: 0.21 Alerting philosophy (symptoms, pages vs tickets, the multi-window burn-rate table), 0.17 Budget arithmetic.
Four properties, one trade
The SRE workbook judges an alert on:
precision of the times it fired, how many were real problems
recall of the real problems, how many made it fire
detection time how long from the start of the problem to the page
reset time how long it keeps firing after the problem is fixed
Every threshold and window trades these against each other. A short window detects fast and resets fast but fires on noise (low precision). A long window is precise but slow to fire and slow to reset. The multi-window, multi-burn-rate design is the compromise that scores well on all four.
Detection time, computed
For a rule "burn rate over window W above threshold B", and a problem that starts suddenly at a constant error ratio e (with a clean window before it), the error ratio over the window after t minutes is e x t / W (the bad minutes are a growing share of the window). It fires when that crosses B x (1 - SLO):
t = W x B x (1 - SLO) / e (only if e > B x (1 - SLO), otherwise never)
For a 99.9% SLO and the workbook's three rules:
$ awk 'BEGIN {n=split("100 50 10 5 2 1 0.7 0.3", r, " "); printf "%-8s %12s %12s %12s\n", "errors", "fast 14.4x", "slow 6x", "ticket 1x"; for (i=1; i<=n; i++) {e=r[i]/100; f=(e>0.0144)?sprintf("%.1f min", 60*0.0144/e):"never"; s=(e>0.006)?sprintf("%.1f min", 360*0.006/e):"never"; t=(e>0.001)?sprintf("%.1f h", 72*0.001/e):"never"; printf "%-8s %12s %12s %12s\n", r[i] "%", f, s, t}}'
errors fast 14.4x slow 6x ticket 1x
100% 0.9 min 2.2 min 0.1 h
50% 1.7 min 4.3 min 0.1 h
10% 8.6 min 21.6 min 0.7 h
5% 17.3 min 43.2 min 1.4 h
2% 43.2 min 108.0 min 3.6 h
1% never 216.0 min 7.2 h
0.7% never 308.6 min 10.3 h
0.3% never never 24.0 h
(awk again as a calculator; cond ? a : b means "a if the condition holds, otherwise b".) The fast rule's long window is 60 minutes, the slow rule's 360, the ticket's 3 days = 72 hours; the short windows fill faster, so the long window decides.
Read the columns: a total outage pages in under a minute. A 10% error rate pages in under 9 minutes. A 1% leak never trips the fast rule - it burns 10x, not 14.4x - but the slow rule catches it in about 3.5 hours, having spent about 5% of the budget. A 0.3% leak (burn 3) is only a ticket, a day later, because at that rate there are 10 days of budget left: a working-hours problem.
In #4471 the error ratio was about 30%. The formula says 60 x 0.0144 / 0.3 = 2.9 minutes; the page came about 4 minutes after the first error, because the first two minutes ran at only 18%.
Reset time, and why two windows
Once the problem is fixed, the long window still remembers it. A 1-hour window that saw 30 minutes of 30% errors stays above 14.4x for most of the next hour: the alert keeps paging after the fix, people learn to ignore "stale" pages, and the next real one is ignored with them.
The short window (1/12 of the long: 5m for 1h, 30m for 6h) fixes that: it drops below the threshold within minutes of the fix, and the rule requires both. In the burn mission (0.23) you will see it happen: at the resolve time the 1-hour burn is higher than at the page - and the alert resolves anyway, because the 5-minute burn is zero.
"for:" is not free
Alert rules can say for: 5m: the condition must stay true for five minutes in a row before the alert fires. That removes flapping, and adds 5 minutes to detection for every incident. With burn-rate windows the long window already provides the stability, so SLO alerts usually use a short for: or none. The cause-based alerts you will meet in the rules mission (CPU > 80% for 5m) need for: precisely because a single reading means nothing.
Low-traffic services
Burn rates are ratios; with little traffic, one request is a big ratio:
$ awk 'BEGIN {for (n=10; n<=100000; n*=10) printf "%6d req/h: 1 error = %.4f%% = burn %.1fx\n", n, 100/n, (1/n)/0.001}'
10 req/h: 1 error = 10.0000% = burn 100.0x
100 req/h: 1 error = 1.0000% = burn 10.0x
1000 req/h: 1 error = 0.1000% = burn 1.0x
10000 req/h: 1 error = 0.0100% = burn 0.1x
100000 req/h: 1 error = 0.0010% = burn 0.0x
At 10 requests an hour a single failure is a 100x burn and pages instantly - and is also, over a month, 7200 requests of which the budget allows 7. Options, from the workbook:
- generate traffic: synthetic probes that run the real journey every minute, so the service always has requests to divide by
- combine small services into one SLO for the journey they serve
- require a minimum count of failed requests (
and errors > 5) before paging - lower the SLO or lengthen the window, if the service truly cannot be measured precisely - do not pretend it can
Dashboards before pages
An alert is only actionable if the person paged can see what it saw. For every burn-rate alert, the linked dashboard should show the burn rate over the alert's own windows (1h and 5m, 6h and 30m), the error ratio, and the budget remaining. Incident #4502 (a mission later in this chapter) is what happens without it: a 6-hour alert judged by a 5-minute graph, and snoozed.
Grouping and inhibition
When the fast and the slow rule both fire for the same outage, nobody needs two pages. The part of a monitoring system that sends the notifications usually groups alerts that share the same labels (tags like service=checkout) into one notification, and inhibits (mutes) one alert while another is firing: slow burn is muted while fast burn fires for the same service. The same mechanism stops one dead server from paging once for every service that runs on it. Duplicate pages are one of the main causes of alert fatigue.
Measuring alert quality
Precision is measurable from the pager export (the list of every page sent, with times), after the fact:
precision = pages where a human had to act / all pages
noise = pages that resolved themselves, or were acked and ignored
load = pages per 12-hour shift, and how many at night
The SRE book's ceiling is about two incidents per 12-hour shift: handling one properly, including the follow-up, takes hours. Review the pages every week at the on-call handover (when one on-call person passes the phone to the next). For each alert that paged: was it actionable? If not, fix the alert - raise the threshold, move it to a ticket, delete it, or automate the response - never "tell the on-call to ignore it". The fatigue mission does this review on a month of real-looking pages.
Recall is harder: it is measured against incidents that were found some other way (customers, support, a colleague). Every postmortem should ask "which alert should have fired, and why didn't it?" - one incident you will review later, #4488, was detected by support after 19 minutes, which means that service's recall was zero.
The checklist for a page
symptom it measures what users experience (an SLO burn, a failing journey)
urgent it will hurt users before working hours if nobody acts
actionable a human can do something; the alert says what service and SLO,
and links a runbook and a dashboard
judgement the response is not a script (if it is, automate it and do not page)
unique nothing else is already paging for the same problem
Anything that fails one of these is a ticket, a dashboard, or nothing. (A script is a file of commands a machine can run in one go - if the response fits in one, nobody needs waking.)
In short
detection t = W x B x (1 - SLO) / e; fast for big outages, slow for leaks
reset the short window (W/12) resolves the alert minutes after the fix
for: adds its duration to every detection
low traffic synthetic traffic, combine services, minimum counts
noise group, inhibit, review every page weekly, fix the alert not the human
What you can now do:
- compute how long a burn-rate rule takes to fire for a given error rate
- explain what the short window and
for:do to detection and reset - measure an alert's precision from a pager export