OnCallReady

Lesson 0.21 · SRE Fundamentals · 11 min read

Alerting philosophy, and multi-window burn rates

In plain words

A smoke alarm that goes off every time you make toast gets its battery taken out, and then it doesn't warn you about a real fire. A good alarm goes off for real danger, early enough to act, and stops once the danger is gone. And it rings about the thing you care about, smoke in the house, not about "the toaster is warm", which may or may not matter.

Good alerting is the same. Page on symptoms users feel, measured as SLO burn rate, not on causes like CPU. Every page must be urgent, actionable and need human judgement. The SRE workbook's multi-window alert pages when the burn rate is above 14.4 over both the last hour and the last 5 minutes: the hour proves it's significant, the 5 minutes lets the alert clear soon after the fix.

Why this lesson

Checkout's team had three alerts. In one day they woke someone up three times for nothing - and when checkout really broke, none of them fired. The person on-call (the engineer whose phone rings when something breaks, taking turns week by week) learned the only lesson such alerts teach: pages are noise. This lesson is about alerts that fire when users hurt, and only then.

What you need to know already: 0.16 Error budgets (burn rate, 14.4), 0.5 Saturation and USE (causes vs the resource that runs out).

Symptoms, not causes

A symptom is what users experience: errors, slowness, "checkout does not work". A cause is a reason it might happen: high CPU, a restarted copy of the service, a disk at 70%, a full connection pool.

Symptoms win, for two reasons that pull in the same direction:

Causes still matter: they are what you look at on the dashboard once the symptom has paged you, and some deserve a ticket.

Rewrite example. Cause: "CPU above 80% for 5 minutes on a checkout copy" - fires during every batch job, and did not fire during an outage where CPU sat at 18%. Symptom: "checkout is burning its 99.9% error budget more than 14.4 times too fast, over both the last hour and the last 5 minutes".

Actionable

A page is an alert that interrupts a human right now - a phone call or push notification, day or night. The SRE book's test for a page. Before an alert pages anybody:

In practice an actionable page also carries a link to a runbook (a written step-by-step guide for handling one kind of problem) or a dashboard, and names the service and the SLO it is protecting.

You're paged at 3am. For that to have been justified: users were being hurt (or about to be), it would not have fixed itself before morning, it needed a human decision, and you were the right human. If any of those is false, the alert gets fixed the next day - not the responder.

Alert fatigue

Alert fatigue is what happens to people who get too many useless alerts. It is caused by: cause-based alerts, thresholds with no link to user impact, flapping alerts (firing and clearing over and over as a number wobbles around the line), duplicates (five alerts for one problem), alerts nobody owns and nobody tunes.

What it costs: people learn that pages are usually noise, so they ack (acknowledge: press "I've got it") without reading, and the one real page is handled slowly or missed. Sleep and morale go, then people. The SRE book's rule of thumb is at most two incidents per 12-hour on-call shift, because handling one properly (including the follow-up) takes hours.

Where each signal goes

WhereWho, whenExamples
Pagea human, nowSLO burn rate fast enough to empty the budget soon; total outage
Ticketa human, within daysslow budget burn; disk predicted to fill in 5 days; certificate expiring in 14 days
Dashboard / logsnobody, until someone is diagnosingCPU, memory, pool usage, restarts, queue length

Multi-window, multi-burn-rate alerts

The SRE workbook's recommended SLO alert (chapter 5, Alerting on SLOs). For a 99.9% SLO:

severity   budget spent   long window   short window   burn rate
page       2%             1 hour        5 minutes      14.4
page       5%             6 hours       30 minutes     6
ticket     10%            3 days        6 hours        1

Each row is one alert rule: "if the burn rate measured over the long window AND the burn rate measured over the short window are both above this number, send this kind of alert". Severity is how serious an alert is, which decides where it goes (page or ticket).

Where 14.4 comes from: spending 2% of a 30-day budget in 1 hour means burning at 0.02 x 720 h / 1 h = 14.4 times the sustainable rate. As an error ratio: 14.4 x 0.001 = 1.44%.

Why two windows. The long window gives significance: it only fires when a meaningful chunk of budget is gone, not on a 30-second blip. The short window gives a fast reset: once the problem is fixed, the 5-minute ratio drops to zero within 5 minutes, and the alert resolves (stops firing) - instead of the 1-hour window keeping it firing for up to an hour after the fix. The alert fires only when both are over the threshold.

Written as a rule for a monitoring tool, the fast one reads like this - "the share of 5xx among checkout's requests over the last 1h is above 14.4 x 0.1%, and the same over the last 5m":

(
  sum(rate(http_requests_total{job="checkout",code=~"5.."}[1h]))
  / sum(rate(http_requests_total{job="checkout"}[1h]))
) > (14.4 * 0.001)
and
(
  sum(rate(http_requests_total{job="checkout",code=~"5.."}[5m]))
  / sum(rate(http_requests_total{job="checkout"}[5m]))
) > (14.4 * 0.001)

How to read it without knowing the language: http_requests_total is a counter of requests the monitoring tool keeps; {job="checkout",code=~"5.."} narrows it to checkout's requests whose status code matches 5.. (5 followed by any two characters: every 5xx); [1h] is the window; rate(...) and sum(...) turn the counter into requests per second; the / divides failures by all requests. You will write this line into a rules file in a mission soon - copy its shape, you do not need to know the language yet.

Later (Ch 28): you will write and test these rules for real in Prometheus, the monitoring tool this syntax (PromQL) belongs to.

The four properties the workbook judges an alert on: precision (does it fire only for real problems), recall (does it fire for every real problem), detection time and reset time. Multi-window burn rates are the compromise that scores well on all four.

What you can now do:

Why it helps

You'll be on call, and the quality of the alerts decides your sleep and whether real incidents are caught. Situations: you inherit "CPU above 80% for 5 minutes", which fired during every batch job and stayed silent during an outage at 18% CPU; you rewrite it as a burn-rate alert. A team is getting ten pages a night and acking without reading; you apply the actionability test and move most to tickets or dashboards. You design the page/ticket/dashboard split for a new service. Interviewers ask "how would you reduce alert fatigue?" and "what is a multi-window burn-rate alert?" often, and a precise answer with the 14.4 derivation stands out.

FAQ

Where does the 14.4 come from?

It is the burn rate that spends 2% of a 30-day budget in one hour: 0.02 times 720 hours divided by 1 hour is 14.4. As an error ratio against a 99.9% SLO, that's 14.4 times 0.001, 1.44% errors. The slower page, 6 over 6 hours, spends 5% of the budget; the ticket, 1 over 3 days, spends 10%.

Why does the alert need two windows?

The long window gives significance: it fires only when a meaningful share of budget is gone, not on a 30-second blip. But it also remembers the problem long after it's fixed, which would keep the alert firing for up to an hour. The short window, a twelfth of the long one, drops to zero minutes after the fix, so requiring both resets the alert quickly.

Should cause-based alerts be deleted entirely?

Mostly they move rather than disappear. Causes like CPU, restarts, memory and pool usage go on the dashboard linked from the symptom alert, where they help diagnosis. Some become tickets, particularly predictions like "disk full in five days" or "certificate expires in 14 days". Only rarely should a cause page, when it predicts user impact before morning.

What makes a page actionable?

It detects something urgent and user-visible now or soon, that nothing else detects; it can't be safely ignored; a human can do something about it, and it needs judgement (a robotic response should be automated); and it carries the service, the SLO it protects, a runbook link and a dashboard link. If any of these fails, fix the alert, not the responder.

How many pages per shift is too many?

The SRE book's rule of thumb is at most two incidents per 12-hour on-call shift, because handling one properly, including follow-up, takes hours. Beyond that, people start acknowledging without reading, the real page gets missed, and sleep and morale suffer. Review page volume and precision weekly in the handover.

In an interview Junior

Why should you alert on symptoms rather than causes?

Because pages should fire when users are hurt, and only then.

Most causes do not hurt anyone: CPU at 85% during a batch job, a restart nobody noticed, a disk at 71% with weeks of room. Paging on them wakes people for nothing, and people learn to ignore pages (alert fatigue). And you cannot list every cause in advance, but almost all of them show up as the same symptom: users getting errors or slow answers. Checkout's outage happened at 18% CPU, so a CPU alert would have stayed silent.

So the page is a symptom, ideally the SLO's burn rate: "checkout is burning its 99.9% budget more than 14.4 times too fast over the last hour and the last 5 minutes". Causes go on the dashboard the page links to, and slow problems become tickets, like "disk full in four days".

Also asked: What makes an alert actionable? · When should something be a page, a ticket, or only a dashboard? · Why does a burn-rate alert check two windows?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.