OnCallReady

Lesson 0.22 · SRE Fundamentals · 17 min read

Designing burn-rate alerts: detection, reset, low traffic, noise

In plain words

Think of a smoke alarm again, with a dial for sensitivity. Turn it up and it catches a tiny wisp of smoke instantly, but also every piece of toast. Turn it down and it never bothers you about toast, but a slow smouldering fire might take ages to set it off. And once the smoke clears, you want it to stop beeping quickly, not keep going for an hour.

Burn-rate alerts have the same dials. Detection time for a rule over window W at threshold B is W x B x (1 - SLO) / e: a total outage pages in under a minute, a 1% leak never trips the fast rule but the slow one catches it in about 3.5 hours. The short window makes reset fast. A "must hold for N minutes" setting adds delay. Low-traffic services need synthetic traffic or minimum counts, and grouping related alerts stops duplicate pages.

Why this lesson

A burn-rate alert fires at 02:10. The on-call opens the dashboard, sees a flat line over the last five minutes, decides it is a glitch, and snoozes it. Six hours later a quarter of the month's budget is gone. The alert was right; the way it was designed and shown made it unbelievable. This lesson is about tuning an alert so people trust it: how fast it fires, how fast it clears, and how noisy it is.

What you need to know already: 0.21 Alerting philosophy (symptoms, pages vs tickets, the multi-window burn-rate table), 0.17 Budget arithmetic.

Four properties, one trade

The SRE workbook judges an alert on:

precision        of the times it fired, how many were real problems
recall           of the real problems, how many made it fire
detection time   how long from the start of the problem to the page
reset time       how long it keeps firing after the problem is fixed

Every threshold and window trades these against each other. A short window detects fast and resets fast but fires on noise (low precision). A long window is precise but slow to fire and slow to reset. The multi-window, multi-burn-rate design is the compromise that scores well on all four.

Detection time, computed

For a rule "burn rate over window W above threshold B", and a problem that starts suddenly at a constant error ratio e (with a clean window before it), the error ratio over the window after t minutes is e x t / W (the bad minutes are a growing share of the window). It fires when that crosses B x (1 - SLO):

t = W x B x (1 - SLO) / e              (only if e > B x (1 - SLO), otherwise never)

For a 99.9% SLO and the workbook's three rules:

$ awk 'BEGIN {n=split("100 50 10 5 2 1 0.7 0.3", r, " "); printf "%-8s %12s %12s %12s\n", "errors", "fast 14.4x", "slow 6x", "ticket 1x"; for (i=1; i<=n; i++) {e=r[i]/100; f=(e>0.0144)?sprintf("%.1f min", 60*0.0144/e):"never"; s=(e>0.006)?sprintf("%.1f min", 360*0.006/e):"never"; t=(e>0.001)?sprintf("%.1f h", 72*0.001/e):"never"; printf "%-8s %12s %12s %12s\n", r[i] "%", f, s, t}}'
errors     fast 14.4x      slow 6x    ticket 1x
100%          0.9 min      2.2 min        0.1 h
50%           1.7 min      4.3 min        0.1 h
10%           8.6 min     21.6 min        0.7 h
5%           17.3 min     43.2 min        1.4 h
2%           43.2 min    108.0 min        3.6 h
1%              never    216.0 min        7.2 h
0.7%            never    308.6 min       10.3 h
0.3%            never        never       24.0 h

(awk again as a calculator; cond ? a : b means "a if the condition holds, otherwise b".) The fast rule's long window is 60 minutes, the slow rule's 360, the ticket's 3 days = 72 hours; the short windows fill faster, so the long window decides.

Read the columns: a total outage pages in under a minute. A 10% error rate pages in under 9 minutes. A 1% leak never trips the fast rule - it burns 10x, not 14.4x - but the slow rule catches it in about 3.5 hours, having spent about 5% of the budget. A 0.3% leak (burn 3) is only a ticket, a day later, because at that rate there are 10 days of budget left: a working-hours problem.

In #4471 the error ratio was about 30%. The formula says 60 x 0.0144 / 0.3 = 2.9 minutes; the page came about 4 minutes after the first error, because the first two minutes ran at only 18%.

Reset time, and why two windows

Once the problem is fixed, the long window still remembers it. A 1-hour window that saw 30 minutes of 30% errors stays above 14.4x for most of the next hour: the alert keeps paging after the fix, people learn to ignore "stale" pages, and the next real one is ignored with them.

The short window (1/12 of the long: 5m for 1h, 30m for 6h) fixes that: it drops below the threshold within minutes of the fix, and the rule requires both. In the burn mission (0.23) you will see it happen: at the resolve time the 1-hour burn is higher than at the page - and the alert resolves anyway, because the 5-minute burn is zero.

"for:" is not free

Alert rules can say for: 5m: the condition must stay true for five minutes in a row before the alert fires. That removes flapping, and adds 5 minutes to detection for every incident. With burn-rate windows the long window already provides the stability, so SLO alerts usually use a short for: or none. The cause-based alerts you will meet in the rules mission (CPU > 80% for 5m) need for: precisely because a single reading means nothing.

Low-traffic services

Burn rates are ratios; with little traffic, one request is a big ratio:

$ awk 'BEGIN {for (n=10; n<=100000; n*=10) printf "%6d req/h: 1 error = %.4f%% = burn %.1fx\n", n, 100/n, (1/n)/0.001}'
    10 req/h: 1 error = 10.0000% = burn 100.0x
   100 req/h: 1 error = 1.0000% = burn 10.0x
  1000 req/h: 1 error = 0.1000% = burn 1.0x
 10000 req/h: 1 error = 0.0100% = burn 0.1x
100000 req/h: 1 error = 0.0010% = burn 0.0x

At 10 requests an hour a single failure is a 100x burn and pages instantly - and is also, over a month, 7200 requests of which the budget allows 7. Options, from the workbook:

Dashboards before pages

An alert is only actionable if the person paged can see what it saw. For every burn-rate alert, the linked dashboard should show the burn rate over the alert's own windows (1h and 5m, 6h and 30m), the error ratio, and the budget remaining. Incident #4502 (a mission later in this chapter) is what happens without it: a 6-hour alert judged by a 5-minute graph, and snoozed.

Grouping and inhibition

When the fast and the slow rule both fire for the same outage, nobody needs two pages. The part of a monitoring system that sends the notifications usually groups alerts that share the same labels (tags like service=checkout) into one notification, and inhibits (mutes) one alert while another is firing: slow burn is muted while fast burn fires for the same service. The same mechanism stops one dead server from paging once for every service that runs on it. Duplicate pages are one of the main causes of alert fatigue.

Measuring alert quality

Precision is measurable from the pager export (the list of every page sent, with times), after the fact:

precision   = pages where a human had to act / all pages
noise       = pages that resolved themselves, or were acked and ignored
load        = pages per 12-hour shift, and how many at night

The SRE book's ceiling is about two incidents per 12-hour shift: handling one properly, including the follow-up, takes hours. Review the pages every week at the on-call handover (when one on-call person passes the phone to the next). For each alert that paged: was it actionable? If not, fix the alert - raise the threshold, move it to a ticket, delete it, or automate the response - never "tell the on-call to ignore it". The fatigue mission does this review on a month of real-looking pages.

Recall is harder: it is measured against incidents that were found some other way (customers, support, a colleague). Every postmortem should ask "which alert should have fired, and why didn't it?" - one incident you will review later, #4488, was detected by support after 19 minutes, which means that service's recall was zero.

The checklist for a page

symptom      it measures what users experience (an SLO burn, a failing journey)
urgent       it will hurt users before working hours if nobody acts
actionable   a human can do something; the alert says what service and SLO,
             and links a runbook and a dashboard
judgement    the response is not a script (if it is, automate it and do not page)
unique       nothing else is already paging for the same problem

Anything that fails one of these is a ticket, a dashboard, or nothing. (A script is a file of commands a machine can run in one go - if the response fits in one, nobody needs waking.)

In short

detection   t = W x B x (1 - SLO) / e; fast for big outages, slow for leaks
reset       the short window (W/12) resolves the alert minutes after the fix
for:        adds its duration to every detection
low traffic synthetic traffic, combine services, minimum counts
noise       group, inhibit, review every page weekly, fix the alert not the human

What you can now do:

Why it helps

When a burn-rate alert pages you, you need to know whether its timing makes sense, and when you design alerts for a service, you need to predict it. Situations: a team asks why a 1% error leak took three hours to page; the formula shows the fast rule can't fire under 1.44%, and that's by design. An alert keeps paging after the fix; the short window is missing. A low-traffic internal tool pages on a single failed request; you add synthetic probes or an errors > 5 condition. #4502 was a 6-hour alert judged against a 5-minute graph and snoozed; the dashboard must show the alert's own windows. This is the level of alert design senior SRE interviews probe.

Commands in this lesson

awk

FAQ

Why did a 1% error rate never trigger the fast-burn page?

The fast rule fires at burn rate 14.4, an error ratio of 1.44% against a 99.9% SLO. At 1% the burn rate is 10, so the fast rule never fires, however long it lasts. The slow rule, burn 6 over 6 hours, catches it in about 3.6 hours, after about 5% of the budget is spent. That's intended: at burn 10 there are three days of budget, so a few hours' detection is acceptable.

Should SLO alerts wait a few minutes before firing?

Usually briefly or not at all. Most alerting tools have a setting that makes the condition hold for, say, five minutes before firing; it removes flapping (firing and clearing over and over) but adds five minutes to every detection. With burn-rate windows, the long window already provides stability. The waiting setting was necessary for cause alerts like "CPU above 80%", where a single sample means nothing.

What is the difference between grouping and inhibition in alerting?

Grouping combines alerts that share a label, like the service name, into one notification, so a server outage doesn't page once for every service on it. Inhibition mutes one alert while another related alert is firing: for example, the slow-burn page is muted while the fast-burn page fires for the same service. Both reduce duplicate pages, a major cause of alert fatigue.

How do I measure recall of my alerts?

Against incidents found some other way: by customers, support, a colleague, or a dashboard someone happened to look at. Every postmortem should ask which alert should have fired and why it didn't. If support detected an outage after 19 minutes, as in #4488, that service's alert recall for that failure was zero, and fixing it is an action item.

What should the dashboard linked from a burn-rate alert show?

The burn rate over the alert's own windows (1 hour and 5 minutes for the fast rule, 6 hours and 30 minutes for the slow one), the error ratio, the budget remaining, and saturation and dependency panels below for diagnosis. A responder judging a 6-hour alert against a 5-minute graph will see nothing and snooze a real problem.

In an interview Junior

A burn-rate alert keeps firing for an hour after the outage was fixed. Why, and how do you fix it?

The rule probably looks only at a long window. A 1-hour error ratio still contains the outage's minutes for up to an hour after the fix, so the burn rate stays above the threshold while users are already fine. People get used to stale pages and start ignoring them.

The fix is the multi-window rule: pair the long window with a short one, 1/12 of its length (5 minutes for 1 hour, 30 minutes for 6 hours), and fire only when both are above the burn rate (14.4 for the fast page). After the fix the 5-minute ratio drops to zero within five minutes, the and turns false and the alert resolves. The long window still makes sure it only fires for real budget loss. I'd also check for a long for: on the rule, which delays both firing and clearing.

Also asked: How long does a burn-rate rule take to fire for a 100% outage, and for a 2% error rate? · How do you alert for a service that only gets a few requests an hour? · How would you measure whether an alert is any good?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.