OnCallReady

Lesson 0.16 · SRE Fundamentals · 10 min read

Error budgets

In plain words

Imagine your parents give you a monthly allowance for treats. You can spend it however you like: sweets, a comic, a cinema ticket. If you spend it all in the first week, no more treats until next month, and you spend the rest of the month without. If you never spend any, maybe the allowance was bigger than you needed.

An error budget is that allowance for failures. With a 99.9% SLO, 0.1% of requests may fail: 861 of 861,501 checkout requests in 30 days, or about 43 minutes of full outage. Deploys, migrations, experiments and incidents all spend it. The burn rate says how fast you are spending: burn rate 1 lasts exactly 30 days, 14.4 spends 2% of the month in one hour. The error budget policy says what happens when it's gone.

Why this lesson

The developers want to release the new payment page on Friday. The people who run the service say "not this week, it has been shaky". Neither side has a number, so the most senior person in the room decides, and everyone leaves annoyed. An error budget replaces that argument with arithmetic both sides agreed to in advance.

What you need to know already: 0.8 SLIs, SLOs and SLAs (good/valid, the 30-day window, what each nine allows).

The budget is the SLO, turned around

error budget  =  1 - SLO
99.9% SLO     =  0.1% of valid requests may fail

The error budget is the amount of failure the SLO allows. Two ways to express it, both from the same number:

request-based   0.1% x 861,501 requests in 30 days  =  861 failed requests
time-based      0.1% x 30 days x 24 h x 60 min       =  43.2 minutes of full outage

(An outage is a period when the service does not work at all.) Request-based is what you actually measure (it weighs a failure at the busiest hour more than one at 4am, when few people are shopping). Time-based is how you explain it to a product owner (the person who decides what the product should do and in what order).

Spending it

The budget is not a failure allowance you hope not to use. It is the amount of risk you are allowed to take. Things that spend it:

A budget that is never spent means the SLO is too loose or the team is too cautious - you could have shipped faster. A budget spent in a week means something has to change.

Why it ends the dev-vs-ops argument

Without a budget: developers ("dev", who build features) want to ship, operations ("ops", who keep things running) want stability, and whoever argues louder or is more senior wins. Every release is a negotiation about feelings.

With a budget, agreed in advance by product, dev and SRE, it is arithmetic:

budget left       ship. take risks. the reliability is already paid for.
budget exhausted  feature releases stop; reliability work, fixes and
                  security patches only, until the service is back in SLO

Nobody has to be the villain. The SLO is a product decision (how reliable do our users need this), and the budget makes the consequence automatic. Developers get a clear, earned licence to move fast; SRE gets a lever that is not "because I said so". It also balances itself: teams that ship carefully get to ship more.

Exhausted halfway through the month

What should happen is whatever the error budget policy says (a short written agreement: what happens at which level of spending) - which is why you write the policy before you need it. A typical one:

  1. Freeze feature launches (no new features go out). P0 fixes (priority 0: the most urgent class of bug), security fixes and reliability work still ship.
  2. The team's priority becomes the reliability problem: postmortems (written reviews of what happened - a later lesson) for the incidents that burned it, and their follow-up tasks, ahead of features.
  3. The freeze lasts until the service is back inside its SLO over the trailing window (for a rolling 30 days: until enough bad days drop out of it).
  4. If people disagree about applying it, it goes up to a named person - it is not renegotiated in the middle of an incident.

It is not a punishment, and nobody is blamed for it. It is the brakes working.

Burn rate

Burn rate is how fast you are spending the budget, compared with spending it exactly evenly over the window:

burn rate  =  observed error ratio  /  (1 - SLO)

burn rate 1     the budget lasts exactly 30 days
burn rate 2     gone in 15 days
burn rate 14.4  2% of the monthly budget gone in one hour
burn rate 720   gone in one hour (30 days = 720 hours)

time to empty  =  720 hours / burn rate

The error ratio is failed requests / all valid requests over some stretch of time. An error ratio of 1.44% sounds small. Against a 99.9% SLO it is a burn rate of 1.44 / 0.1 = 14.4 - and that is exactly the level at which the alerting lesson's page fires.

What you can now do:

Why it helps

The error budget is what ends the "developers want to ship, ops wants stability" argument, and you will be the one explaining it to a product owner. Situations: the budget is exhausted on the 15th and a product manager wants the big launch anyway; the pre-agreed policy (freeze features, reliability work first) settles it without anyone being the villain. The budget has been untouched for three months; you argue the team can ship faster or take on the risky migration in daylight. An alert says a 1.44% error rate, which sounds small; as a burn rate of 14.4 it's a page. Interviewers ask "what happens when the error budget runs out?" to see whether you know it's a policy, not a punishment.

FAQ

Should the error budget be measured in minutes or requests?

Enforce it in requests, which is what you actually measure and which weighs a failure at peak more than one at 4am. Explain it in minutes, because "43 minutes of full outage per month" is easy for a product owner to understand. Both come from the same number: 1 minus the SLO.

Is an unspent error budget good news?

Not necessarily. A budget that is never spent means the SLO is looser than the service's real reliability, or the team is more cautious than it needs to be. The budget is permission to take risk: ship faster, run migrations, do chaos experiments. Consistently leaving most of it unused suggests you could move faster, or tighten the SLO.

What happens when the error budget is exhausted?

Whatever the error budget policy, agreed in advance, says. Typically feature launches freeze while P0 fixes, security fixes and reliability work continue; postmortem action items for the incidents that burned it take priority; the freeze lasts until the service is back inside its SLO over the trailing window; disagreements escalate to a named person rather than being renegotiated mid-incident.

Does a dependency's outage count against my budget?

Yes, if your users were affected. Users don't care whose fault it was, and the SLO describes their experience. Repeated budget loss to a dependency is a signal to add resilience (timeouts, fallbacks, caching, running in more than one location) or to renegotiate what you can promise. It's a reliability problem to solve, not an exception to carve out.

What is a burn rate?

The speed of budget spending relative to spending it exactly over the window: the observed error ratio divided by 1 minus the SLO. Burn rate 1 empties the budget in exactly 30 days, 2 in 15 days, 720 in one hour. Time to empty is 720 hours divided by the burn rate. It lets you compare the same error ratio against different SLOs.

In an interview Junior

What is an error budget and how is it used?

The error budget is 1 - SLO: how much failure the service is allowed over the window. For 99.9% over 30 days that is 0.1% of valid requests, or about 43 minutes of full outage.

Everything risky spends it: incidents, bad releases, maintenance, experiments. While budget is left, the team ships; the reliability users need is already paid for. When it runs out, an error budget policy agreed in advance applies - typically feature releases pause and the team works on reliability until the service is back within its SLO. That turns the old "dev wants to ship, ops wants stability" argument into arithmetic.

How fast it is spent is the burn rate: error ratio / budget fraction. A 1% error rate against a 99.9% SLO is a burn rate of 10 - the monthly budget gone in three days.

Also asked: How much downtime does a 99.9% SLO allow per month? · What is a burn rate, and what does a burn rate of 1 mean? · Who decides what happens when the error budget runs out?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.