Why this lesson
Checkout's team had three alerts. In one day they woke someone up three times for nothing - and when checkout really broke, none of them fired. The person on-call (the engineer whose phone rings when something breaks, taking turns week by week) learned the only lesson such alerts teach: pages are noise. This lesson is about alerts that fire when users hurt, and only then.
What you need to know already: 0.16 Error budgets (burn rate, 14.4), 0.5 Saturation and USE (causes vs the resource that runs out).
Symptoms, not causes
A symptom is what users experience: errors, slowness, "checkout does not work". A cause is a reason it might happen: high CPU, a restarted copy of the service, a disk at 70%, a full connection pool.
Symptoms win, for two reasons that pull in the same direction:
- Precision (of the times the alert fired, how many were real problems). Most causes do not hurt anyone. CPU at 85% during a batch job, one copy of a service restarting while the others carry the traffic, a disk at 71% with weeks of headroom - each is a page at 3am for nothing.
- Recall (of the real problems, how many made the alert fire). You cannot list every cause in advance. A slow database query, an expired certificate (the file that proves a site's identity for HTTPS; it has an end date), a configuration change that shrank the pool - you did not write an alert for it, but it still produces the symptom. A symptom alert catches causes you have never imagined.
Causes still matter: they are what you look at on the dashboard once the symptom has paged you, and some deserve a ticket.
Rewrite example. Cause: "CPU above 80% for 5 minutes on a checkout copy" - fires during every batch job, and did not fire during an outage where CPU sat at 18%. Symptom: "checkout is burning its 99.9% error budget more than 14.4 times too fast, over both the last hour and the last 5 minutes".
Actionable
A page is an alert that interrupts a human right now - a phone call or push notification, day or night. The SRE book's test for a page. Before an alert pages anybody:
- Does it detect something urgent, actionable and user-visible now or soon, that nothing else detects?
- Could I ever ignore it, knowing it is harmless? Then it should not page.
- Can I do something? Is it urgent, or could it wait until morning?
- Does responding require judgement? If the response is robotic ("restart it"), a machine should do it and nobody should be woken.
In practice an actionable page also carries a link to a runbook (a written step-by-step guide for handling one kind of problem) or a dashboard, and names the service and the SLO it is protecting.
You're paged at 3am. For that to have been justified: users were being hurt (or about to be), it would not have fixed itself before morning, it needed a human decision, and you were the right human. If any of those is false, the alert gets fixed the next day - not the responder.
Alert fatigue
Alert fatigue is what happens to people who get too many useless alerts. It is caused by: cause-based alerts, thresholds with no link to user impact, flapping alerts (firing and clearing over and over as a number wobbles around the line), duplicates (five alerts for one problem), alerts nobody owns and nobody tunes.
What it costs: people learn that pages are usually noise, so they ack (acknowledge: press "I've got it") without reading, and the one real page is handled slowly or missed. Sleep and morale go, then people. The SRE book's rule of thumb is at most two incidents per 12-hour on-call shift, because handling one properly (including the follow-up) takes hours.
Where each signal goes
| Where | Who, when | Examples |
|---|---|---|
| Page | a human, now | SLO burn rate fast enough to empty the budget soon; total outage |
| Ticket | a human, within days | slow budget burn; disk predicted to fill in 5 days; certificate expiring in 14 days |
| Dashboard / logs | nobody, until someone is diagnosing | CPU, memory, pool usage, restarts, queue length |
Multi-window, multi-burn-rate alerts
The SRE workbook's recommended SLO alert (chapter 5, Alerting on SLOs). For a 99.9% SLO:
severity budget spent long window short window burn rate
page 2% 1 hour 5 minutes 14.4
page 5% 6 hours 30 minutes 6
ticket 10% 3 days 6 hours 1
Each row is one alert rule: "if the burn rate measured over the long window AND the burn rate measured over the short window are both above this number, send this kind of alert". Severity is how serious an alert is, which decides where it goes (page or ticket).
Where 14.4 comes from: spending 2% of a 30-day budget in 1 hour means burning at 0.02 x 720 h / 1 h = 14.4 times the sustainable rate. As an error ratio: 14.4 x 0.001 = 1.44%.
Why two windows. The long window gives significance: it only fires when a meaningful chunk of budget is gone, not on a 30-second blip. The short window gives a fast reset: once the problem is fixed, the 5-minute ratio drops to zero within 5 minutes, and the alert resolves (stops firing) - instead of the 1-hour window keeping it firing for up to an hour after the fix. The alert fires only when both are over the threshold.
Written as a rule for a monitoring tool, the fast one reads like this - "the share of 5xx among checkout's requests over the last 1h is above 14.4 x 0.1%, and the same over the last 5m":
(
sum(rate(http_requests_total{job="checkout",code=~"5.."}[1h]))
/ sum(rate(http_requests_total{job="checkout"}[1h]))
) > (14.4 * 0.001)
and
(
sum(rate(http_requests_total{job="checkout",code=~"5.."}[5m]))
/ sum(rate(http_requests_total{job="checkout"}[5m]))
) > (14.4 * 0.001)
How to read it without knowing the language: http_requests_total is a counter of requests the monitoring tool keeps; {job="checkout",code=~"5.."} narrows it to checkout's requests whose status code matches 5.. (5 followed by any two characters: every 5xx); [1h] is the window; rate(...) and sum(...) turn the counter into requests per second; the / divides failures by all requests. You will write this line into a rules file in a mission soon - copy its shape, you do not need to know the language yet.
Later (Ch 28): you will write and test these rules for real in Prometheus, the monitoring tool this syntax (PromQL) belongs to.
The four properties the workbook judges an alert on: precision (does it fire only for real problems), recall (does it fire for every real problem), detection time and reset time. Multi-window burn rates are the compromise that scores well on all four.
What you can now do:
- rewrite a cause-based alert as a symptom-based one
- decide whether something should be a page, a ticket or a dashboard
- explain why a burn-rate alert uses two windows