Why this lesson
"How much budget do we have left, and when does it run out?" is the question a product owner will ask you on the morning after an incident. "Some" is not an answer. This lesson is the arithmetic behind a real one: "we have spent 107%; at today's rate we would have been out by Thursday anyway".
What you need to know already: 0.16 Error budgets (1 - SLO, burn rate), 0.8 SLIs (rolling windows).
One number, several units
The budget is always 1 - SLO, in whatever unit the SLI counts. In minutes of total outage, for the windows you will meet:
$ awk 'BEGIN {split("99 99.5 99.9 99.95 99.99", s, " "); printf "%-7s %10s %10s %10s %10s\n", "SLO", "week", "28d", "30d", "quarter"; for (i=1; i<=5; i++) {b=1-s[i]/100; printf "%-7s %9.1fm %9.1fm %9.1fm %9.1fm\n", s[i], b*7*1440, b*28*1440, b*30*1440, b*91*1440}}'
SLO week 28d 30d quarter
99 100.8m 403.2m 432.0m 1310.4m
99.5 50.4m 201.6m 216.0m 655.2m
99.9 10.1m 40.3m 43.2m 131.0m
99.95 5.0m 20.2m 21.6m 65.5m
99.99 1.0m 4.0m 4.3m 13.1m
(awk as a calculator again: split cuts the text "99 99.5 ..." into a list s, the for walks it, and printf lines the columns up. 1440 is minutes per day.) Memorise one cell - 99.9% over 30 days is 43.2 minutes - and derive the rest: each extra nine divides by ten, 99.5% is five times 99.9%.
In requests, the budget depends on traffic. From checkout's 30-day export (a TSV file: tab-separated values, one row per day - date, valid requests, 5xx):
$ cd ~/oncall-lab/labs/0-sre
$ awk -F'\t' '!/^#/ {r+=$2; e+=$3; n++} END {b=r*0.001; printf "days %d requests %d errors %d budget %.1f consumed %.1f%%\n", n, r, e, b, 100*e/b}' data/checkout-30d.tsv
days 30 requests 861501 errors 922 budget 861.5 consumed 107.0%
-F'\t' splits on tabs only; !/^#/ skips the comment lines at the top; r+=$2 adds up column 2. 861 failed requests is the whole month's allowance, and 922 have failed.
Partial outages
Time-based budgets assume a total outage. Real incidents are partial - some requests fail, most do not. Convert with the error ratio:
equivalent full-outage minutes = duration x error ratio
37 minutes at ~30% errors = 11 minutes of full outage
4 hours at 1% errors = 2.4 minutes
2 minutes at 100% = 2 minutes
#4471 (checkout's outage from 0.5) failed 216 of the 729 requests in its 37-minute window, about 30%: roughly 11 of the 43 budget minutes, or 26%. Request-based said 25% - close here, because checkout's traffic is flat. They drift apart when an incident hits the daily peak (more requests per minute than average: the request-based share is higher) or the middle of the night (lower). Enforce in requests; explain in minutes.
Where the budget went
$ sort -t$'\t' -k3,3nr data/checkout-30d.tsv | head -3
2026-09-11 29004 368
2026-09-22 24040 223
2026-09-05 29424 14
sort flags: -t$'\t' makes the tab the column separator ($'\t' is how bash writes a tab character), -k3,3nr sorts on column 3 only, as numbers (n), largest first (r = reverse). head -3 keeps the top three.
Two days hold most of it: the failover on the 11th (368 errors, 43% of the budget) and today (223, 26%). The other 28 days are background noise of about 12 errors a day, together 38%. That breakdown is what the reliability plan should follow: which kinds of event eat the budget. Here a failover procedure and the release process each cost more than everything else combined.
Burn rate
burn rate = error ratio over some window / (1 - SLO)
The speed of spending, relative to spending the budget exactly over the SLO window. For a 99.9% SLO:
$ awk 'BEGIN {for (i=1; i<=6; i++) {split("0.05 0.1 0.5 1 1.44 5", r, " "); e=r[i]/100; b=e/0.001; printf "error ratio %5.2f%% burn %6.1fx 30d budget gone in %7.1f h\n", r[i], b, 720/b}}'
error ratio 0.05% burn 0.5x 30d budget gone in 1440.0 h
error ratio 0.10% burn 1.0x 30d budget gone in 720.0 h
error ratio 0.50% burn 5.0x 30d budget gone in 144.0 h
error ratio 1.00% burn 10.0x 30d budget gone in 72.0 h
error ratio 1.44% burn 14.4x 30d budget gone in 50.0 h
error ratio 5.00% burn 50.0x 30d budget gone in 14.4 h
720 is the hours in 30 days. A "1% error rate", which sounds small, empties a 99.9% budget in three days. A 0.05% error rate is burn 0.5: the month ends with half the budget unspent.
The same error ratio means different things under different SLOs: 1% is burn 10 against 99.9%, and burn 2 against 99.5%. Convert to burn rate before deciding how bad something is.
Time left, not time total
The table assumes a full budget. Mid-month, with part of it spent:
hours left = (1 - consumed) x window_hours / burn rate
consumed 60%, burning at 6x, 30-day window: 0.40 x 720 / 6 = 48 hours
consumed 95%, burning at 2x: 0.05 x 720 / 2 = 18 hours
consumed 107%: 0 - it is gone, whatever the burn
This is the number to put in front of a product owner: "at the current error rate we are out of budget on Thursday afternoon".
Latency has a budget too
A latency SLO of "99% of requests under 300 ms" has a 1% budget of slow requests, with its own consumption and burn rate. Inside #4471's window about 98% of requests were slower than 300 ms: the latency budget burned at about 98x (0.98 / 0.01) while availability burned at about 300x (0.30 / 0.001). Separate SLIs, separate budgets, separate alerts.
Spending it on purpose
A budget that is always nearly full is not an achievement - it means the SLO is looser than the service's real reliability, or releases are more cautious than they need to be. With budget left, a team can:
- ship faster (fewer approval steps, bigger releases)
- run the risky migration now, in daylight, with the budget as the safety margin
- run chaos experiments and game days (planned practice sessions where the team breaks something on purpose and rehearses the response) that deliberately spend some of it
The error budget policy says what happens at the other end: for example 50% consumed in the first week triggers a review, 100% freezes features. Checkout's policy (data/error-budget-policy.md) is the minimal version.
Rolling windows recover slowly
With a rolling window the budget "comes back" only as bad days age out: the 368 errors of the 11th count against checkout until they are 30 days old. A second incident during that time extends the freeze. The freeze mission computes the exact day.
In short
budget 1 - SLO; 99.9% over 30d = 43.2 min = 0.1% of requests
partial minutes x error ratio; enforce in requests
burn rate error ratio / (1 - SLO); 1 = exactly on budget
time to empty window / burn; time LEFT = (1 - consumed) x window / burn
recovery only as bad events leave the rolling window
What you can now do:
- compute a budget in minutes and requests for any SLO and window
- find which days or events ate the budget
- say how many hours are left at the current burn rate