Why this lesson
"Is checkout reliable?" Ask five people and you get five answers: "it's fine", "it was down on Tuesday", "the dashboard is green", "customers are complaining". None of them is a number, so nobody can agree on whether to stop and fix things or keep shipping. SLIs, SLOs and SLAs turn "reliable" into a measurement, a target and a promise - three different things with three different owners.
What you need to know already: 0.1 The four golden signals (requests, 5xx, latency), 0.2 Percentiles.
Three words, three different jobs
SLI indicator a MEASUREMENT "96.0% of requests succeeded"
SLO objective a TARGET for it "99.9% of requests succeed, over 30 days"
SLA agreement a CONTRACT "below 99.5% in a month, you get 10% credit"
SLI = Service Level Indicator, SLO = Service Level Objective, SLA = Service Level Agreement. "Service level" just means "how well the service is doing".
SLI - a measurement
The SRE workbook's form (Google's second SRE book, the practical companion to the first), and the one to use: good events / valid events, as a percentage. An event is one thing that happened that you can count - here, one request.
availability SLI = requests that did not return 5xx / valid requests
latency SLI = requests served in <= 300 ms / valid requests
The ratio form matters: 0% is all bad, 100% is all good, and every SLI reads the same way on a dashboard.
The two decisions that make or break an SLI:
- What is "valid". Health checks (automatic "are you alive?" requests from monitoring tools - also called probes) are not users. They hit a cheap address that almost never fails, so counting them inflates the SLI. Most teams also leave 4xx out of "bad" (the client's mistake), and decide explicitly about 429 (Too Many Requests: you throttled them) and 499 (nginx's code for "the client gave up and hung up" - often because you were slow).
- Where it is measured. Measured at the load balancer or edge proxy (the first machine that receives users' requests and spreads them across the copies of the service - nginx on this box), it sees what users see, including requests the service never answered. Measured inside the service, it misses every request that died before reaching it.
SLI types
| Type | Question | Example |
|---|---|---|
| Availability | did it answer successfully? | non-5xx / valid |
| Latency | did it answer fast enough? | requests under 300 ms / valid |
| Correctness (quality) | was the answer right? | orders whose total matches the cart / orders checked |
Correctness is the one people skip and the one that hurts most: a service that returns 200 with the wrong price is 100% available and completely broken. It usually needs a prober (a small program that makes real requests on a schedule and checks the answers) or a reconciliation job (a program that later compares records, e.g. charged amounts against carts) to measure. Data pipelines and storage add freshness (is the data recent?), coverage (was everything processed?) and durability (is stored data still there?).
SLO - a target
An SLO is an SLI, a target and a window (the stretch of time you measure over): 99.9% of valid checkout requests return a non-5xx response, measured over a rolling 30 days. "Rolling" means "always the last 30 days, recomputed as time moves on".
Picking a number that is not arbitrary:
- Start from users. What failure rate or delay do they notice, complain about, or abandon a cart over? That sets the minimum.
- Look at history. What has the service actually delivered? It tells you what is achievable - but do not simply lock in today's performance, or you promise reliability you only had by luck.
- Check dependencies (the other services yours needs in order to work). You cannot be more available than something you hard-depend on. Dependencies in a chain multiply: two 99.9% services, one calling the other, give ~99.8%.
- Price each nine. "Nines" is how people say availability: 99.9% is "three nines". Every extra nine costs roughly ten times more, and shrinks the room you have to change things.
- Keep it loose at first, tighten when you have data. Few SLOs, simple ones.
What each nine buys, over 30 days
99% 7.2 hours of failure allowed
99.5% 3.6 hours
99.9% 43.2 minutes
99.95% 21.6 minutes
99.99% 4.3 minutes
99.999% 26 seconds
(30 days = 43,200 minutes; 0.1% of that is 43.2.)
Why never 100%
- Users cannot tell. Their phone, Wi-Fi and internet provider fail far more than 0.01% of the time. The difference between 99.99% and 100% is invisible to them.
- It is infinitely expensive. Each nine costs roughly 10x the last; the last one never arrives.
- It forbids change. 100% leaves zero room for failure: no releases, no migrations (moving data or a system to a new setup), no experiments - because every change is a risk.
- Your dependencies are not 100% either, so the target is a lie from day one.
SLA - a contract
An SLA is a promise to a customer with a consequence: service credits (money off the bill), refunds, the right to cancel the contract. The easy test, from the SRE book: what happens if it is missed? If the answer is "money or a contract clause", it is an SLA. If the answer is "we stop shipping features and fix things", it is an SLO.
Commercially, the SLA is set looser than the SLO (say 99.5% vs 99.9%), so the internal alarm goes off - and the team reacts - long before a breach costs money. Engineers own SLOs; legal and sales own SLAs. An SLA with no SLO under it means you find out you are in breach from the customer's complaint about the invoice.
A web API, in order
An API (application programming interface) is a service's set of addresses that other programs call - here everything under /api/.
SLI % of valid HTTP requests to /api/* (probes excluded), measured at the edge,
that return a non-5xx status; and % that complete in <= 300 ms
SLO 99.9% availability and 99% of requests under 300 ms, rolling 30 days
SLA 99.5% monthly availability; below that, 10% service credit
What you can now do:
- define SLI, SLO and SLA for a web API, in that order
- say what counts as a valid event and why probes are left out
- explain why 100% is the wrong target