OnCallReady

Lesson 0.17 · SRE Fundamentals · 16 min read

Budget arithmetic: minutes, requests, burn rate, time left

In plain words

Say you get 30 sweets for the month, one a day on average. If you eat 2 a day, they last 15 days. If you already ate 18 of them and you're eating 3 a day, you have 12 left, so 4 more days. And if you ate half a sweet, that counts as half: a partial treat is a partial spend.

Error budget arithmetic is exactly that. 99.9% over 30 days is 43.2 minutes, each extra nine divides by ten. A partial outage counts as duration times error ratio: 37 minutes at 30% errors is about 11 minutes. Burn rate is error ratio divided by 1 minus the SLO, and time left is (1 - consumed) x window hours / burn rate. Budgets on a rolling window only come back as bad days age out.

Why this lesson

"How much budget do we have left, and when does it run out?" is the question a product owner will ask you on the morning after an incident. "Some" is not an answer. This lesson is the arithmetic behind a real one: "we have spent 107%; at today's rate we would have been out by Thursday anyway".

What you need to know already: 0.16 Error budgets (1 - SLO, burn rate), 0.8 SLIs (rolling windows).

One number, several units

The budget is always 1 - SLO, in whatever unit the SLI counts. In minutes of total outage, for the windows you will meet:

$ awk 'BEGIN {split("99 99.5 99.9 99.95 99.99", s, " "); printf "%-7s %10s %10s %10s %10s\n", "SLO", "week", "28d", "30d", "quarter"; for (i=1; i<=5; i++) {b=1-s[i]/100; printf "%-7s %9.1fm %9.1fm %9.1fm %9.1fm\n", s[i], b*7*1440, b*28*1440, b*30*1440, b*91*1440}}'
SLO           week        28d        30d    quarter
99          100.8m     403.2m     432.0m    1310.4m
99.5         50.4m     201.6m     216.0m     655.2m
99.9         10.1m      40.3m      43.2m     131.0m
99.95         5.0m      20.2m      21.6m      65.5m
99.99         1.0m       4.0m       4.3m      13.1m

(awk as a calculator again: split cuts the text "99 99.5 ..." into a list s, the for walks it, and printf lines the columns up. 1440 is minutes per day.) Memorise one cell - 99.9% over 30 days is 43.2 minutes - and derive the rest: each extra nine divides by ten, 99.5% is five times 99.9%.

In requests, the budget depends on traffic. From checkout's 30-day export (a TSV file: tab-separated values, one row per day - date, valid requests, 5xx):

$ cd ~/oncall-lab/labs/0-sre
$ awk -F'\t' '!/^#/ {r+=$2; e+=$3; n++} END {b=r*0.001; printf "days %d  requests %d  errors %d  budget %.1f  consumed %.1f%%\n", n, r, e, b, 100*e/b}' data/checkout-30d.tsv
days 30  requests 861501  errors 922  budget 861.5  consumed 107.0%

-F'\t' splits on tabs only; !/^#/ skips the comment lines at the top; r+=$2 adds up column 2. 861 failed requests is the whole month's allowance, and 922 have failed.

Partial outages

Time-based budgets assume a total outage. Real incidents are partial - some requests fail, most do not. Convert with the error ratio:

equivalent full-outage minutes = duration x error ratio

37 minutes at ~30% errors      = 11 minutes of full outage
4 hours at 1% errors           = 2.4 minutes
2 minutes at 100%              = 2 minutes

#4471 (checkout's outage from 0.5) failed 216 of the 729 requests in its 37-minute window, about 30%: roughly 11 of the 43 budget minutes, or 26%. Request-based said 25% - close here, because checkout's traffic is flat. They drift apart when an incident hits the daily peak (more requests per minute than average: the request-based share is higher) or the middle of the night (lower). Enforce in requests; explain in minutes.

Where the budget went

$ sort -t$'\t' -k3,3nr data/checkout-30d.tsv | head -3
2026-09-11	29004	368
2026-09-22	24040	223
2026-09-05	29424	14

sort flags: -t$'\t' makes the tab the column separator ($'\t' is how bash writes a tab character), -k3,3nr sorts on column 3 only, as numbers (n), largest first (r = reverse). head -3 keeps the top three.

Two days hold most of it: the failover on the 11th (368 errors, 43% of the budget) and today (223, 26%). The other 28 days are background noise of about 12 errors a day, together 38%. That breakdown is what the reliability plan should follow: which kinds of event eat the budget. Here a failover procedure and the release process each cost more than everything else combined.

Burn rate

burn rate  =  error ratio over some window  /  (1 - SLO)

The speed of spending, relative to spending the budget exactly over the SLO window. For a 99.9% SLO:

$ awk 'BEGIN {for (i=1; i<=6; i++) {split("0.05 0.1 0.5 1 1.44 5", r, " "); e=r[i]/100; b=e/0.001; printf "error ratio %5.2f%%  burn %6.1fx  30d budget gone in %7.1f h\n", r[i], b, 720/b}}'
error ratio  0.05%  burn    0.5x  30d budget gone in  1440.0 h
error ratio  0.10%  burn    1.0x  30d budget gone in   720.0 h
error ratio  0.50%  burn    5.0x  30d budget gone in   144.0 h
error ratio  1.00%  burn   10.0x  30d budget gone in    72.0 h
error ratio  1.44%  burn   14.4x  30d budget gone in    50.0 h
error ratio  5.00%  burn   50.0x  30d budget gone in    14.4 h

720 is the hours in 30 days. A "1% error rate", which sounds small, empties a 99.9% budget in three days. A 0.05% error rate is burn 0.5: the month ends with half the budget unspent.

The same error ratio means different things under different SLOs: 1% is burn 10 against 99.9%, and burn 2 against 99.5%. Convert to burn rate before deciding how bad something is.

Time left, not time total

The table assumes a full budget. Mid-month, with part of it spent:

hours left = (1 - consumed) x window_hours / burn rate

consumed 60%, burning at 6x, 30-day window:   0.40 x 720 / 6   = 48 hours
consumed 95%, burning at 2x:                  0.05 x 720 / 2   = 18 hours
consumed 107%:                                0 - it is gone, whatever the burn

This is the number to put in front of a product owner: "at the current error rate we are out of budget on Thursday afternoon".

Latency has a budget too

A latency SLO of "99% of requests under 300 ms" has a 1% budget of slow requests, with its own consumption and burn rate. Inside #4471's window about 98% of requests were slower than 300 ms: the latency budget burned at about 98x (0.98 / 0.01) while availability burned at about 300x (0.30 / 0.001). Separate SLIs, separate budgets, separate alerts.

Spending it on purpose

A budget that is always nearly full is not an achievement - it means the SLO is looser than the service's real reliability, or releases are more cautious than they need to be. With budget left, a team can:

The error budget policy says what happens at the other end: for example 50% consumed in the first week triggers a review, 100% freezes features. Checkout's policy (data/error-budget-policy.md) is the minimal version.

Rolling windows recover slowly

With a rolling window the budget "comes back" only as bad days age out: the 368 errors of the 11th count against checkout until they are 30 days old. A second incident during that time extends the freeze. The freeze mission computes the exact day.

In short

budget          1 - SLO; 99.9% over 30d = 43.2 min = 0.1% of requests
partial         minutes x error ratio; enforce in requests
burn rate       error ratio / (1 - SLO); 1 = exactly on budget
time to empty   window / burn; time LEFT = (1 - consumed) x window / burn
recovery        only as bad events leave the rolling window

What you can now do:

Why it helps

These numbers are what you say out loud during and after incidents. "At the current error rate we're out of budget on Thursday afternoon" is the sentence that gets a product owner's attention. Situations: a postmortem needs the impact as a percentage of budget; the request-based 25% and the minutes-based 26% for #4471 are close only because traffic was flat. A dashboard says 1% errors; you convert: burn 10 against 99.9%, burn 2 against 99.5%, very different urgency. Planning a freeze, you compute the day the 11th's 368 errors age out of the window. Interviews often include one of these calculations, and doing it quickly and correctly is convincing.

Commands in this lesson

awk cd sort

FAQ

What is the fastest way to remember the budget in minutes?

Memorise one cell: 99.9% over 30 days is 43.2 minutes. Each extra nine divides by ten (99.99% is 4.3 minutes, 99.999% about 26 seconds), each nine fewer multiplies by ten (99% is 7.2 hours), and 99.5% is five times 99.9%, 3.6 hours. A week is 7/30 of those numbers; 28 days is 28/30.

How do I count a partial outage against a time-based budget?

Multiply duration by the error ratio: 37 minutes at about 30% errors is roughly 11 minutes of full outage; 4 hours at 1% is 2.4 minutes. That converts real incidents, which are rarely total, into the time-based budget. For enforcement, count failed requests directly, which also accounts for whether the incident hit peak traffic or the middle of the night.

Why does 1% error rate mean different things for different services?

Because urgency depends on the budget, not the raw rate. Burn rate is error ratio over 1 minus the SLO: 1% is burn 10 against a 99.9% SLO, emptying the month in three days, and burn 2 against 99.5%, emptying it in fifteen. Always convert to burn rate before deciding how bad something is.

Does a latency SLO have its own budget?

Yes. "99% of requests under 300 ms" has a 1% budget of slow requests, with its own consumption, burn rate and alerts. During #4471, about 98% of requests were slower than 300 ms, so the latency budget burned at about 98x, while availability burned at about 300x. Separate SLIs, separate budgets.

How long does it take to recover budget on a rolling window?

Budget comes back only as the bad events leave the window. With a rolling 30 days, the 368 errors from the 11th count until they're 30 days old, regardless of how well the service behaves in between. A second incident during that time extends any freeze. You can compute the exact day by summing daily errors in the trailing window going forward.

In an interview Junior

How much downtime does a 99.9% SLO allow per month?

About 43 minutes over 30 days: 30 x 24 x 60 = 43,200 minutes, and 0.1% of that is 43.2. Scale from there: 99% is ten times more (7.2 hours), 99.99% ten times less (about 4.3 minutes).

Two things I'd add. Most incidents are partial, so convert with duration x error ratio: 37 minutes at 30% errors uses about 11 of the 43 minutes. And the budget is usually counted in requests, not minutes: 0.1% of the month's valid requests, so an hour of errors at peak traffic costs more than an hour at 4am, which matches what users felt. To see how long is left, divide the budget remaining by the current burn rate - time left, not the 30 days total.

Also asked: A 20-minute incident failed 25% of requests. How much of a 99.9% monthly budget did it use? · How do you find which days or events used up the error budget? · Why does a rolling 30-day window take so long to recover after a big incident?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.