From chapter 0 to a rule file
The threshold alert from the first incident (28.13) paged for nothing and slept through a real outage, and nobody could say why 5% was the number. Chapter 0 gave you a better idea - alert on how fast the error budget is being spent - but worked it out on paper. This lesson turns it into Prometheus rules that page the right amount.
What you need to know already: SLI, SLO, error budget (0.8, 0.9, 0.16); burn rate and budget arithmetic (0.17); multi-window burn-rate alerts, detection time and reset time as ideas (0.21, 0.22); rate(), sum by and increase() (27.8); buckets and le (27.15); and / or between vectors (27.15, 28.1); recording rules (27.27); promtool tests (28.4).
Quick reminder of the Ch 0 words: an SLI is good events / valid events; an SLO is the target for that ratio over a window (99.9% over 30 days); the error budget is 1 minus the SLO (0.1% of requests may fail); the burn rate is how fast you are spending it. This is the Google SRE Workbook's multi-window, multi-burn-rate design (the Workbook is Google's free book of SRE practice), with the arithmetic done out loud.
The SLI in PromQL
For orders: the proportion of requests to the API that do not fail with a server error, over 30 days, target 99.9%.
bad = sum(rate(http_server_requests_seconds_count{job="orders", status=~"5..", uri!="/actuator/health"}[W]))
valid = sum(rate(http_server_requests_seconds_count{job="orders", uri!="/actuator/health"}[W]))
error ratio over W = bad / valid
W stands for a window you fill in (5m, 1h...). http_server_requests_seconds_count is the orders request counter (21.11, 27.15); status=~"5.." keeps statuses that are "5 followed by any two characters"; uri!="/actuator/health" drops health checks.
Decisions hidden in that query, each of which you must be able to defend:
- 5xx only. A 404 or a 400 is the client's fault and does not burn the budget. A 429 (rate-limited) is debatable; decide and write it down.
- Health checks excluded. Kubernetes probes (17.20) are cheap requests that almost never fail; including them dilutes the ratio and hides user-facing errors.
- Measured at the server. Requests that never reach the app (the load balancer is down) are invisible here. Measuring at nginx or the load balancer catches more; synthetic probes (0.8, 0.9) catch the rest.
A latency SLI uses a bucket: "requests served in 250 ms or less" is rate(..._bucket{le="0.25"}[W]) over rate(..._count[W]). It only works if a bucket boundary sits exactly at the threshold (27.15).
Burn rate
burn rate = error ratio over a window / (1 - SLO)
With a 99.9% SLO the budget is 0.1%. An error ratio of 0.1% is burn rate 1: at that pace the budget runs out exactly at the end of the 30 days. An error ratio of 1.44% is burn rate 14.4.
How much budget does a burn consume? burn rate x (window / SLO period). A 30-day period is 720 hours, so:
burn 14.4 for 1 hour = 14.4 x 1/720 = 2% of the monthly budget
burn 6 for 6 hours = 6 x 6/720 = 5%
burn 3 for 1 day = 3 x 24/720 = 10%
burn 1 for 3 days = 1 x 72/720 = 10%
Those four lines are the Workbook's recommended thresholds, read backwards: page if you have burnt 2% of the month's budget in an hour, or 5% in six hours; open a ticket at 10% in a day or 10% in three days. 14.4 is not magic - it is "2% of a 30-day budget in one hour".
Why one window is not enough
A single long window is slow and sticky. Alert on "burn > 14.4 over the last hour" and a total outage (100% errors) crosses 1.44% after 52 seconds - fine. But after the outage is fixed, the 1-hour window stays above 1.44% for up to an hour, so the alert keeps firing long after users are fine, and a second incident in that hour is invisible.
A single short window is noisy. "Burn > 14.4 over 5 minutes" fires on a five-minute blip that consumed 0.17% of the budget - nobody should wake for that.
Both together: fire only when the long window says "this is significant" and the short window says "and it is still happening". The long window gives precision; the short window gives a fast reset once the problem stops. The Workbook uses a short window of 1/12 of the long one:
severity long window short window burn rate budget consumed
page 1h 5m 14.4 2%
page 6h 30m 6 5%
ticket 1d 2h 3 10%
ticket 3d 6h 1 10%
Each row is one condition: "the long-window burn rate is above the number, and so is the short-window one". The two page rows become one alert, the two ticket rows another.
Detection time
How long until the page, for a given outage? For an error ratio that starts suddenly, after elapsed minutes the 1h window's ratio is error ratio x elapsed / window (the rest of the hour had no errors), and it pages when that reaches the threshold:
99.9% SLO, 1h/14.4 page, threshold ratio 1.44%
outage error ratio time to page
100% 1h x 1.44/100 = 0.9 min
50% 1.7 min
10% 8.6 min
5% 17 min
2% 43 min
1.44% or less never, from this rule - the 6h rule and the tickets catch it
That table is the honest answer to "how fast does your alerting detect an outage": proportional to how bad it is, by design.
The rules
Recording rules first, one per window, so the alerts stay readable and every window is computed once:
groups:
- name: orders-slo
rules:
- record: job:slo_errors_per_request:ratio_rate5m
expr: |
sum by (job) (rate(http_server_requests_seconds_count{job="orders", status=~"5..", uri!="/actuator/health"}[5m]))
/
sum by (job) (rate(http_server_requests_seconds_count{job="orders", uri!="/actuator/health"}[5m]))
# ... the same for 30m, 1h, 2h, 6h, 1d, 3d
- alert: OrdersErrorBudgetBurn
expr: |
(
job:slo_errors_per_request:ratio_rate1h > (14.4 * 0.001)
and
job:slo_errors_per_request:ratio_rate5m > (14.4 * 0.001)
)
or
(
job:slo_errors_per_request:ratio_rate6h > (6 * 0.001)
and
job:slo_errors_per_request:ratio_rate30m > (6 * 0.001)
)
labels:
severity: page
annotations:
summary: "orders is burning its 30-day error budget too fast"
- alert: OrdersErrorBudgetBurnSlow
expr: |
(
job:slo_errors_per_request:ratio_rate1d > (3 * 0.001)
and
job:slo_errors_per_request:ratio_rate2h > (3 * 0.001)
)
or
(
job:slo_errors_per_request:ratio_rate3d > (1 * 0.001)
and
job:slo_errors_per_request:ratio_rate6h > (1 * 0.001)
)
labels:
severity: ticket
The recording-rule names follow the level:metric:operations convention from 27.27: aggregated by job, the metric "SLO errors per request", computed as a ratio of rates over 5 minutes. and keeps a series only if the other side has one with the same labels; or returns the series of either side - so the page alert fires if either page row of the table is true.
Why this shape:
sum by (job)on both sides so the ratio has exactly thejoblabel andandmatches the windows one to one. Mismatched labels between windows is the classic way these rules silently never fire.14.4 * 0.001is written out, not pre-multiplied to0.0144: the reader sees the burn rate and the budget separately.- no
for: the short window already requires the problem to be current, and aforwould only delay the page. (Afor: 2mis sometimes added to ride out one bad scrape; it costs two minutes of detection.) [3d]over raw counters is expensive at every evaluation; in real deployments the long windows are computed from the short recorded ratio (avg_over_time(job:...:ratio_rate5m[3d])- which is an approximation because it averages ratios instead of summing requests; accurate enough under steady traffic) or evaluated in a rule group with a slowerinterval(a group can set its own evaluation interval).
Error budget remaining
A dashboard number, not an alert:
1 - (
sum(increase(http_server_requests_seconds_count{job="orders", status=~"5..", uri!="/actuator/health"}[30d]))
/
sum(increase(http_server_requests_seconds_count{job="orders", uri!="/actuator/health"}[30d]))
) / 0.001
increase(...[30d]) counts the requests of the last 30 days (27.8), so the inner division is the month's error ratio, and dividing it by the budget (0.001) says what fraction of the budget is used. 1 minus that: 1 = untouched, 0 = spent, negative = SLO breached. It needs 30 days of data - on this box, with six hours, it reads optimistic.
Caveats that come up in interviews
- Low traffic. At 1 request per minute, one error is a 20% error ratio over 5 minutes (1 of 5) - 200x the budget of a 99.9% SLO. The maths still holds, but the alert fires on single requests. Fix: longer windows, a minimum-traffic clause (
and sum(rate(...[1h])) > 1, "and there is at least one request per second"), or synthetic traffic. - The SLO is per user journey, not per endpoint - but the alert should say which endpoint is burning. Add
urito the aggregation of a second, debug rule, or put the breakdown on the dashboard linked from the annotation. - Test it. An SLO alert that never fired is an untested alert: write the promtool test with a synthetic outage, and then fire it for real with a load test (this chapter does both, 28.15 and 28.22).
What you can now do
- Write an SLI as a ratio of rates and defend what it counts and excludes.
- Turn an SLO into the four burn-rate conditions and their recording rules.
- Say how long a given outage takes to page, and why two windows per condition.