OnCallReady

Lesson 0.8 · SRE Fundamentals · 11 min read

SLIs, SLOs and SLAs

In plain words

Imagine a school canteen promising parents good lunches. First they decide what to count: "lunches served hot and on time, out of all lunches ordered" (the indicator). Then they set a goal: "at least 99 out of 100, over each month" (the objective). Separately, the council's contract says that if they fall below 95, the canteen pays a fine (the agreement). The goal is stricter than the contract, so the cooks notice problems long before any fine.

That is SLI, SLO and SLA. The SLI is good events over valid events, like non-5xx checkout requests over valid requests measured at the edge, probes excluded. The SLO is a target and a window, 99.9% over a rolling 30 days. The SLA is a customer contract with a consequence, set looser, like 99.5%.

Why this lesson

"Is checkout reliable?" Ask five people and you get five answers: "it's fine", "it was down on Tuesday", "the dashboard is green", "customers are complaining". None of them is a number, so nobody can agree on whether to stop and fix things or keep shipping. SLIs, SLOs and SLAs turn "reliable" into a measurement, a target and a promise - three different things with three different owners.

What you need to know already: 0.1 The four golden signals (requests, 5xx, latency), 0.2 Percentiles.

Three words, three different jobs

SLI   indicator   a MEASUREMENT      "96.0% of requests succeeded"
SLO   objective   a TARGET for it    "99.9% of requests succeed, over 30 days"
SLA   agreement   a CONTRACT         "below 99.5% in a month, you get 10% credit"

SLI = Service Level Indicator, SLO = Service Level Objective, SLA = Service Level Agreement. "Service level" just means "how well the service is doing".

SLI - a measurement

The SRE workbook's form (Google's second SRE book, the practical companion to the first), and the one to use: good events / valid events, as a percentage. An event is one thing that happened that you can count - here, one request.

availability SLI  =  requests that did not return 5xx  /  valid requests
latency SLI       =  requests served in <= 300 ms      /  valid requests

The ratio form matters: 0% is all bad, 100% is all good, and every SLI reads the same way on a dashboard.

The two decisions that make or break an SLI:

SLI types

TypeQuestionExample
Availabilitydid it answer successfully?non-5xx / valid
Latencydid it answer fast enough?requests under 300 ms / valid
Correctness (quality)was the answer right?orders whose total matches the cart / orders checked

Correctness is the one people skip and the one that hurts most: a service that returns 200 with the wrong price is 100% available and completely broken. It usually needs a prober (a small program that makes real requests on a schedule and checks the answers) or a reconciliation job (a program that later compares records, e.g. charged amounts against carts) to measure. Data pipelines and storage add freshness (is the data recent?), coverage (was everything processed?) and durability (is stored data still there?).

SLO - a target

An SLO is an SLI, a target and a window (the stretch of time you measure over): 99.9% of valid checkout requests return a non-5xx response, measured over a rolling 30 days. "Rolling" means "always the last 30 days, recomputed as time moves on".

Picking a number that is not arbitrary:

  1. Start from users. What failure rate or delay do they notice, complain about, or abandon a cart over? That sets the minimum.
  2. Look at history. What has the service actually delivered? It tells you what is achievable - but do not simply lock in today's performance, or you promise reliability you only had by luck.
  3. Check dependencies (the other services yours needs in order to work). You cannot be more available than something you hard-depend on. Dependencies in a chain multiply: two 99.9% services, one calling the other, give ~99.8%.
  4. Price each nine. "Nines" is how people say availability: 99.9% is "three nines". Every extra nine costs roughly ten times more, and shrinks the room you have to change things.
  5. Keep it loose at first, tighten when you have data. Few SLOs, simple ones.

What each nine buys, over 30 days

99%       7.2 hours of failure allowed
99.5%     3.6 hours
99.9%     43.2 minutes
99.95%    21.6 minutes
99.99%    4.3 minutes
99.999%   26 seconds

(30 days = 43,200 minutes; 0.1% of that is 43.2.)

Why never 100%

SLA - a contract

An SLA is a promise to a customer with a consequence: service credits (money off the bill), refunds, the right to cancel the contract. The easy test, from the SRE book: what happens if it is missed? If the answer is "money or a contract clause", it is an SLA. If the answer is "we stop shipping features and fix things", it is an SLO.

Commercially, the SLA is set looser than the SLO (say 99.5% vs 99.9%), so the internal alarm goes off - and the team reacts - long before a breach costs money. Engineers own SLOs; legal and sales own SLAs. An SLA with no SLO under it means you find out you are in breach from the customer's complaint about the invoice.

A web API, in order

An API (application programming interface) is a service's set of addresses that other programs call - here everything under /api/.

SLI   % of valid HTTP requests to /api/* (probes excluded), measured at the edge,
      that return a non-5xx status; and % that complete in <= 300 ms
SLO   99.9% availability and 99% of requests under 300 ms, rolling 30 days
SLA   99.5% monthly availability; below that, 10% service credit

What you can now do:

Why it helps

Writing an SLO for a service is one of the first real deliverables an SRE is asked for, and one of the most common interview questions: "define an SLO for this API". The traps are in this lesson: counting health checks inflates the SLI; measuring in the app misses every request that died at the proxy; forgetting correctness means a service returning the wrong price looks 100% available. Knowing what each nine costs (99.9% is 43 minutes a month, 99.99% is 4 minutes) lets you push back when someone asks for "five nines" for an internal tool, and knowing why the SLA sits below the SLO shows commercial awareness.

FAQ

Why is 100% never the right SLO?

Users can't tell the difference beyond a point, because their phones, Wi-Fi and ISPs fail far more often. Each extra nine costs roughly ten times more, and the last never arrives. 100% means an error budget of zero, which forbids deploys, migrations and experiments. And your dependencies aren't 100% either, so the target is false from day one.

Should 4xx responses count as errors?

Usually not: the service answered correctly, and the client asked for something wrong (bad input, not found, unauthorised). But decide explicitly for some codes: 429 is bad if you throttled legitimate users because you were overloaded, and nginx's 499 (client closed the connection) can mean users gave up because you were slow. Write the decision down in the SLI definition.

Where should an availability SLI be measured?

Usually at the edge proxy or load balancer: it sees what users see, including requests no server answered, like 502s and 504s from the proxy. Measured inside the application, the SLI misses every request that died before reaching it, and a crashed program reports nothing. Measuring in the user's app or browser is closest to the truth but noisy; synthetic probers (robots acting like users) work at 3am with no traffic but test only one path.

How do I pick the SLO number?

Start from users: what failure rate or delay do they notice or abandon over? Look at history to see what is achievable, without locking in lucky performance. Check hard dependencies, since serial services multiply (two 99.9% services give about 99.8%). Price each extra nine, and start loose with few, simple SLOs, tightening when you have data.

How do I tell an SLA from an SLO?

Ask what happens when it is missed. If the answer is money, credits or a contract clause, it is an SLA, owned by legal and sales. If the answer is "we stop shipping features and fix things", it is an SLO, owned by engineering. The SLA is set looser than the SLO so the internal alarm goes off well before a breach costs money.

In an interview Junior

What is the difference between an SLI, an SLO and an SLA?

A quick test: if missing it costs money, it is an SLA; if it changes what the team works on, it is an SLO. And never aim for 100%: users cannot tell 99.99% from 100%, and it would forbid every change.

Also asked: Give an example of an availability SLI and a latency SLI for a web API. · How much downtime does each extra nine allow over 30 days? · Why is the SLA looser than the SLO?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.