OnCallReady

Lesson 28.14 · Observability II: Alerting, Alertmanager & SLO alerts · 19 min read

SLO alerts in Prometheus: multi-window, multi-burn-rate

In plain words

Imagine you get 100 euros of pocket money per month to spend on snacks. Spending some is fine; that is what it is for. What matters is how fast you spend it. If you spend 2 euros in one afternoon, you would run out in a few days at that pace, so your parents step in now. If you spend a little extra over a whole week, a quiet talk at the weekend is enough.

The SLO is the rule "99.9% of orders requests succeed over 30 days"; the 0.1% that may fail is the pocket money, the error budget. The burn rate is the spending speed. Burn-rate alerts page when 2% of the month's budget goes in an hour, and open a ticket for slow leaks. Two windows, like 1h and 5m, make sure it is significant and still happening.

From chapter 0 to a rule file

The threshold alert from the first incident (28.13) paged for nothing and slept through a real outage, and nobody could say why 5% was the number. Chapter 0 gave you a better idea - alert on how fast the error budget is being spent - but worked it out on paper. This lesson turns it into Prometheus rules that page the right amount.

What you need to know already: SLI, SLO, error budget (0.8, 0.9, 0.16); burn rate and budget arithmetic (0.17); multi-window burn-rate alerts, detection time and reset time as ideas (0.21, 0.22); rate(), sum by and increase() (27.8); buckets and le (27.15); and / or between vectors (27.15, 28.1); recording rules (27.27); promtool tests (28.4).

Quick reminder of the Ch 0 words: an SLI is good events / valid events; an SLO is the target for that ratio over a window (99.9% over 30 days); the error budget is 1 minus the SLO (0.1% of requests may fail); the burn rate is how fast you are spending it. This is the Google SRE Workbook's multi-window, multi-burn-rate design (the Workbook is Google's free book of SRE practice), with the arithmetic done out loud.

The SLI in PromQL

For orders: the proportion of requests to the API that do not fail with a server error, over 30 days, target 99.9%.

bad   = sum(rate(http_server_requests_seconds_count{job="orders", status=~"5..", uri!="/actuator/health"}[W]))
valid = sum(rate(http_server_requests_seconds_count{job="orders", uri!="/actuator/health"}[W]))
error ratio over W = bad / valid

W stands for a window you fill in (5m, 1h...). http_server_requests_seconds_count is the orders request counter (21.11, 27.15); status=~"5.." keeps statuses that are "5 followed by any two characters"; uri!="/actuator/health" drops health checks.

Decisions hidden in that query, each of which you must be able to defend:

A latency SLI uses a bucket: "requests served in 250 ms or less" is rate(..._bucket{le="0.25"}[W]) over rate(..._count[W]). It only works if a bucket boundary sits exactly at the threshold (27.15).

Burn rate

burn rate = error ratio over a window / (1 - SLO)

With a 99.9% SLO the budget is 0.1%. An error ratio of 0.1% is burn rate 1: at that pace the budget runs out exactly at the end of the 30 days. An error ratio of 1.44% is burn rate 14.4.

How much budget does a burn consume? burn rate x (window / SLO period). A 30-day period is 720 hours, so:

burn 14.4 for 1 hour  = 14.4 x 1/720  = 2% of the monthly budget
burn 6    for 6 hours = 6 x 6/720     = 5%
burn 3    for 1 day   = 3 x 24/720    = 10%
burn 1    for 3 days  = 1 x 72/720    = 10%

Those four lines are the Workbook's recommended thresholds, read backwards: page if you have burnt 2% of the month's budget in an hour, or 5% in six hours; open a ticket at 10% in a day or 10% in three days. 14.4 is not magic - it is "2% of a 30-day budget in one hour".

Why one window is not enough

A single long window is slow and sticky. Alert on "burn > 14.4 over the last hour" and a total outage (100% errors) crosses 1.44% after 52 seconds - fine. But after the outage is fixed, the 1-hour window stays above 1.44% for up to an hour, so the alert keeps firing long after users are fine, and a second incident in that hour is invisible.

A single short window is noisy. "Burn > 14.4 over 5 minutes" fires on a five-minute blip that consumed 0.17% of the budget - nobody should wake for that.

Both together: fire only when the long window says "this is significant" and the short window says "and it is still happening". The long window gives precision; the short window gives a fast reset once the problem stops. The Workbook uses a short window of 1/12 of the long one:

severity  long window  short window  burn rate  budget consumed
page      1h           5m            14.4       2%
page      6h           30m           6          5%
ticket    1d           2h            3          10%
ticket    3d           6h            1          10%

Each row is one condition: "the long-window burn rate is above the number, and so is the short-window one". The two page rows become one alert, the two ticket rows another.

Detection time

How long until the page, for a given outage? For an error ratio that starts suddenly, after elapsed minutes the 1h window's ratio is error ratio x elapsed / window (the rest of the hour had no errors), and it pages when that reaches the threshold:

99.9% SLO, 1h/14.4 page, threshold ratio 1.44%

outage error ratio   time to page
100%                 1h x 1.44/100 = 0.9 min
 50%                 1.7 min
 10%                 8.6 min
  5%                 17 min
  2%                 43 min
  1.44% or less      never, from this rule - the 6h rule and the tickets catch it

That table is the honest answer to "how fast does your alerting detect an outage": proportional to how bad it is, by design.

The rules

Recording rules first, one per window, so the alerts stay readable and every window is computed once:

groups:
  - name: orders-slo
    rules:
      - record: job:slo_errors_per_request:ratio_rate5m
        expr: |
          sum by (job) (rate(http_server_requests_seconds_count{job="orders", status=~"5..", uri!="/actuator/health"}[5m]))
            /
          sum by (job) (rate(http_server_requests_seconds_count{job="orders", uri!="/actuator/health"}[5m]))
      # ... the same for 30m, 1h, 2h, 6h, 1d, 3d

      - alert: OrdersErrorBudgetBurn
        expr: |
          (
            job:slo_errors_per_request:ratio_rate1h > (14.4 * 0.001)
            and
            job:slo_errors_per_request:ratio_rate5m > (14.4 * 0.001)
          )
          or
          (
            job:slo_errors_per_request:ratio_rate6h > (6 * 0.001)
            and
            job:slo_errors_per_request:ratio_rate30m > (6 * 0.001)
          )
        labels:
          severity: page
        annotations:
          summary: "orders is burning its 30-day error budget too fast"

      - alert: OrdersErrorBudgetBurnSlow
        expr: |
          (
            job:slo_errors_per_request:ratio_rate1d > (3 * 0.001)
            and
            job:slo_errors_per_request:ratio_rate2h > (3 * 0.001)
          )
          or
          (
            job:slo_errors_per_request:ratio_rate3d > (1 * 0.001)
            and
            job:slo_errors_per_request:ratio_rate6h > (1 * 0.001)
          )
        labels:
          severity: ticket

The recording-rule names follow the level:metric:operations convention from 27.27: aggregated by job, the metric "SLO errors per request", computed as a ratio of rates over 5 minutes. and keeps a series only if the other side has one with the same labels; or returns the series of either side - so the page alert fires if either page row of the table is true.

Why this shape:

Error budget remaining

A dashboard number, not an alert:

1 - (
  sum(increase(http_server_requests_seconds_count{job="orders", status=~"5..", uri!="/actuator/health"}[30d]))
    /
  sum(increase(http_server_requests_seconds_count{job="orders", uri!="/actuator/health"}[30d]))
) / 0.001

increase(...[30d]) counts the requests of the last 30 days (27.8), so the inner division is the month's error ratio, and dividing it by the budget (0.001) says what fraction of the budget is used. 1 minus that: 1 = untouched, 0 = spent, negative = SLO breached. It needs 30 days of data - on this box, with six hours, it reads optimistic.

Caveats that come up in interviews

What you can now do

Why it helps

This is the alerting design senior SRE interviews ask about most, and the one that fixes noisy on-call in practice. When your team replaces twenty threshold alerts with two burn-rate alerts per user journey, pages start correlating with real user pain. You need to explain it convincingly: why 14.4 is "2% of a 30-day budget in an hour", why a single 1h window keeps paging for an hour after recovery, and why a 5% outage pages in about 17 minutes by design.

The SLI decisions matter in real reviews: exclude health checks, count only 5xx, decide about 429, and know that server-side measurement misses a dead load balancer. And the gotchas are practical: mismatched labels between windows make the and never match, low traffic makes one error a 100% ratio, and a [3d] window over raw counters is expensive at every evaluation.

FAQ

What is the difference between SLI, SLO and SLA?

An SLI is a measurement: good events divided by valid events, for example successful requests over all requests. An SLO is the internal target for that SLI over a window, like 99.9% over 30 days. An SLA is a contract with customers that has consequences, such as credits, if breached; it is usually looser than the SLO, so you have room to react before you owe anything.

Where does the number 14.4 come from?

It is the burn rate that consumes 2% of a 30-day error budget in one hour. A 30-day period is 720 hours; budget consumed is burn rate times window over period, so 14.4 times 1/720 equals 0.02. The other Workbook thresholds follow the same way: burn 6 for 6 hours is 5%, burn 3 for a day is 10%, burn 1 for three days is 10%.

Why use two windows per condition?

The long window, such as 1h, decides that enough budget was burnt to matter, so short blips do not page. The short window, about 1/12 of the long one, such as 5m, requires that the problem is still happening. Without it, the 1h ratio stays above threshold for up to an hour after recovery, keeping the alert firing and hiding a second incident. Both must exceed the threshold.

Should health check requests count in the SLI?

Usually not. Kubernetes probes and load balancer checks are frequent, cheap and almost never fail, so they dilute the error ratio and can hide user-facing errors. Exclude them, as with uri!="/actuator/health". Similarly, decide consciously what counts as bad: 5xx yes, 4xx usually no since they are client errors, 429 is a judgement call. Write the decisions down with the SLO.

What goes wrong with burn-rate alerts on low-traffic services?

With very few requests, one failure is a huge ratio over a short window: at one request per minute, a single 5xx in five minutes is a 20% error ratio. The maths is right but the alert fires on individual requests. Options: longer windows, a minimum-traffic condition like and sum(rate(...[1h])) > 1, synthetic traffic so there is always a baseline, or grouping small services into one SLO.

In an interview Mid

Walk me through writing an SLO burn-rate alert for a 99.9% availability SLO.

  1. SLI as a ratio of rates: 5xx over all requests, excluding health checks. Record it per window: job:slo_errors_per_request:ratio_rate5m, ..._rate1h, ..._rate30m, ..._rate6h.
  2. Budget = 1 - 0.999 = 0.1%. Burn rate = error ratio / 0.001; budget used = burn x window / 720 h.
  3. Conditions (the Workbook's table): page when the 1h and 5m ratios are both above 14.4 * 0.001 (2% of the month in an hour) or 6h and 30m above 6 * 0.001 (5%); ticket at 3 (1d/2h) and 1 (3d/6h).
(job:slo_errors_per_request:ratio_rate1h > (14.4 * 0.001)
   and job:slo_errors_per_request:ratio_rate5m > (14.4 * 0.001))
or
(job:slo_errors_per_request:ratio_rate6h > (6 * 0.001)
   and job:slo_errors_per_request:ratio_rate30m > (6 * 0.001))

Why two windows: the long one gives significance (a blip does not page), the short one makes it reset as soon as the problem stops. No for needed. Detection time scales with severity: a total outage pages in under a minute.

Watch: matching labels between windows (sum by (job) on both sides), low traffic (add a minimum-traffic clause), and test it with promtool.

Also asked: What are SLIs, SLOs and error budgets? · Why does each burn-rate condition use two windows? · How do SLO burn-rate alerts behave on a low-traffic service?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.