Why this lesson
The developers want to release the new payment page on Friday. The people who run the service say "not this week, it has been shaky". Neither side has a number, so the most senior person in the room decides, and everyone leaves annoyed. An error budget replaces that argument with arithmetic both sides agreed to in advance.
What you need to know already: 0.8 SLIs, SLOs and SLAs (good/valid, the 30-day window, what each nine allows).
The budget is the SLO, turned around
error budget = 1 - SLO
99.9% SLO = 0.1% of valid requests may fail
The error budget is the amount of failure the SLO allows. Two ways to express it, both from the same number:
request-based 0.1% x 861,501 requests in 30 days = 861 failed requests
time-based 0.1% x 30 days x 24 h x 60 min = 43.2 minutes of full outage
(An outage is a period when the service does not work at all.) Request-based is what you actually measure (it weighs a failure at the busiest hour more than one at 4am, when few people are shopping). Time-based is how you explain it to a product owner (the person who decides what the product should do and in what order).
Spending it
The budget is not a failure allowance you hope not to use. It is the amount of risk you are allowed to take. Things that spend it:
- incidents (unplanned events that hurt users) and outages - the obvious one
- bad releases that get rolled back (replaced again by the previous version)
- planned maintenance that causes errors, migrations (moving data or a system to a new setup), failovers (switching to a backup system)
- experiments: a new database, a risky configuration change, a chaos test (breaking something on purpose to see whether the system copes)
- a dependency's bad day - users do not care whose fault it was
A budget that is never spent means the SLO is too loose or the team is too cautious - you could have shipped faster. A budget spent in a week means something has to change.
Why it ends the dev-vs-ops argument
Without a budget: developers ("dev", who build features) want to ship, operations ("ops", who keep things running) want stability, and whoever argues louder or is more senior wins. Every release is a negotiation about feelings.
With a budget, agreed in advance by product, dev and SRE, it is arithmetic:
budget left ship. take risks. the reliability is already paid for.
budget exhausted feature releases stop; reliability work, fixes and
security patches only, until the service is back in SLO
Nobody has to be the villain. The SLO is a product decision (how reliable do our users need this), and the budget makes the consequence automatic. Developers get a clear, earned licence to move fast; SRE gets a lever that is not "because I said so". It also balances itself: teams that ship carefully get to ship more.
Exhausted halfway through the month
What should happen is whatever the error budget policy says (a short written agreement: what happens at which level of spending) - which is why you write the policy before you need it. A typical one:
- Freeze feature launches (no new features go out). P0 fixes (priority 0: the most urgent class of bug), security fixes and reliability work still ship.
- The team's priority becomes the reliability problem: postmortems (written reviews of what happened - a later lesson) for the incidents that burned it, and their follow-up tasks, ahead of features.
- The freeze lasts until the service is back inside its SLO over the trailing window (for a rolling 30 days: until enough bad days drop out of it).
- If people disagree about applying it, it goes up to a named person - it is not renegotiated in the middle of an incident.
It is not a punishment, and nobody is blamed for it. It is the brakes working.
Burn rate
Burn rate is how fast you are spending the budget, compared with spending it exactly evenly over the window:
burn rate = observed error ratio / (1 - SLO)
burn rate 1 the budget lasts exactly 30 days
burn rate 2 gone in 15 days
burn rate 14.4 2% of the monthly budget gone in one hour
burn rate 720 gone in one hour (30 days = 720 hours)
time to empty = 720 hours / burn rate
The error ratio is failed requests / all valid requests over some stretch of time. An error ratio of 1.44% sounds small. Against a 99.9% SLO it is a burn rate of 1.44 / 0.1 = 14.4 - and that is exactly the level at which the alerting lesson's page fires.
What you can now do:
- turn an SLO into an error budget in minutes and in requests
- say what spends a budget and what an error budget policy decides
- compute a burn rate and how long the budget lasts at that rate