Why this lesson
Two teams measure the same service over the same day. One reports 96.7%, the other 93.7%. Both used "the access log". The difference is three small choices nobody wrote down - and on a 99.9% target, the whole allowance for failure is 0.1 points. This lesson is about making those choices on purpose.
What you need to know already: 0.8 SLIs, SLOs and SLAs (good / valid, probes), 0.2 Percentiles (jq, select, sort | uniq -c).
An SLI is two decisions and a division
SLI = good events / valid events over a window
The SRE workbook separates the SLI specification (what you want to measure, in words: "the proportion of checkout requests that succeed") from the SLI implementation (how you actually count it: "non-5xx responses logged by the edge nginx, excluding /healthz"). Most SLI bugs live in the gap between the two. You will investigate one later in this chapter: the specification meant "searches that succeed", the implementation counted probes and treated a user giving up as a success.
Decision 1: which events are valid
Start from every request the edge saw and remove what is not a user trying to do something:
- probes and monitors: health checks, uptime checkers, the load balancer's own checks, synthetic transactions (scripted fake users that run a real journey, like "log in, add to cart, pay" - they get their own SLI, see below)
- internal traffic that is not part of the user journey (batch jobs - big scheduled tasks - or crawlers you control), if it has different expectations
- sometimes methods or routes outside the SLO's scope (an HTTP method is the verb of a request:
GETreads,POSTsends;OPTIONSis a technical pre-check browsers make;/metricsis where monitoring tools read numbers)
On checkout, the status classes of the valid requests:
$ jq -r 'select(.uri != "/healthz") | .status' /var/log/nginx/checkout.access.log | awk '{c[int($1/100)"xx"]++} END {for (k in c) print k, c[k]}' | sort
2xx 5057
4xx 127
5xx 216
$ jq -r 'select(.uri != "/healthz") | .status' /var/log/nginx/checkout.access.log | sort | uniq -c
4283 200
774 201
40 400
37 401
21 404
29 409
170 500
3 502
43 504
int($1/100)"xx" turns 404 into "4xx": divide by 100, drop the decimals, and glue "xx" on (awk joins values written next to each other). c[...]++ counts one per line under that key; the END loop prints every key and its count.
Decision 2: which events are good
| Status | Usually | Why |
|---|---|---|
| 2xx, 3xx | good | the service did its job |
| 400, 401, 403, 404, 409, 422 | good | the service answered correctly; the client asked for something wrong |
| 429 | depends | good if the client went over a documented limit; bad if you turned away real users because you were overloaded |
| 499 (nginx) | depends | the client hung up first. Bad when your slowness caused it (a phone app's timeout); noise when a search box cancels the previous keystroke's request |
| 5xx | bad | the service failed |
| 200 with an error body | bad | the classic hidden failure - only a correctness SLI or a content check sees it |
The same log, three honest-looking definitions:
good valid SLI
non-5xx over valid (the SLO's definition) 5184 5400 96.00%
non-5xx over every line (probe counted) 6264 6480 96.67%
2xx only over valid (4xx counted as bad) 5057 5400 93.65%
A 3-point spread from definitions alone - on a 99.9% SLO the whole allowance for failure is 0.1 points. That is why the definition is written down, agreed, and reviewed as carefully as code.
Decision 3: where it is measured
client (the app or browser) closest to the user; sees the whole trip over the internet.
Noisy, delayed, and you do not control it.
edge / load balancer sees every request that reached you, including those no
copy of the service answered (502/504 from the proxy).
The usual choice.
application sees only requests that reached the service. Misses every
request that died in front of it - a crashed service
reports nothing at all.
synthetic prober a small program that runs the journey every minute. Works at 3am
with no traffic; measures one path, not your users.
(Measuring in the user's browser is called RUM, real user monitoring.) For checkout the edge nginx is right: it logged the 504s that checkout itself never knew about - checkout did not know nginx had given up on it.
SLI types, with the command that counts them
Availability - you have done it all chapter. Latency - the fraction served fast enough:
$ jq -r 'select(.uri != "/healthz") | .request_time' /var/log/nginx/checkout.access.log | awk '$1 <= 0.3 {f++} END {printf "%d of %d = %.2f%% under 300 ms\n", f, NR, 100*f/NR}'
4651 of 5400 = 86.13% under 300 ms
($1 <= 0.3 {f++}: for every line whose value is at most 0.3 seconds, add one to f.) A latency SLI keeps all valid requests in the bottom of the fraction - a 5 s 504 is a slow request as well as a failed one. Some teams use two thresholds: "90% under 100 ms and 99% under 400 ms" captures both the typical and the tail experience.
Correctness (quality) - was the answer right? The log cannot tell you: a 200 with the wrong price looks perfect. You measure it with a prober that checks the response body, or a reconciliation job ("orders whose charged amount equals the cart total / orders checked"). Freshness for data pipelines ("reads of data newer than 10 minutes / reads"), coverage ("records processed / records received"), durability for storage.
Per route, or per journey
$ jq -r 'select(.uri != "/healthz") | "\(.uri) \(.status)"' /var/log/nginx/checkout.access.log | sed -E 's,/[0-9]+ , ,' | awk '{n[$1]++; if ($2 >= 500) b[$1]++} END {for (k in n) printf "%-26s %5d %4d %7.2f%%\n", k, n[k], b[k], 100*(1-b[k]/n[k])}' | sort
/api/cart 1870 82 95.61%
/api/cart/items 1100 43 96.09%
/api/checkout 817 31 96.21%
/api/orders 815 33 95.95%
/api/payments/authorize 527 19 96.39%
/api/shipping/quote 271 8 97.05%
Columns: route, requests, failures, SLI. Two details. sed -E 's,/[0-9]+ , ,' (sed edits text as it flows past; s,A,B, replaces A with B, and -E enables patterns like [0-9]+ = "one or more digits") turns /api/orders/40117 into /api/orders: without it every order number is its own "route", and you get hundreds of rows with one request each. That is the cardinality problem (too many distinct values of something you group by) - it also breaks monitoring tools, so never group by an id. And "\(.uri) \(.status)" is jq's way of building a text from fields.
Here the incident hit every route equally, so one service-level SLI is enough. When routes differ (a cheap GET /cart vs an expensive POST /checkout), a single SLI lets the busy cheap route hide failures on the important one. Group routes into user journeys ("browse", "pay") and give the important ones their own SLO.
Windows
Rolling (the last 30 days, recomputed continuously) is what users experience and what the SLO itself should use: every bad day counts for exactly 30 days. Calendar (this month) resets on the 1st: an outage on the 30th is forgotten two days later. Calendar windows are for reporting and contracts (SLAs are billed monthly).
28 days is a popular alternative to 30: every window then holds the same number of each weekday, so weekly traffic patterns do not skew one window against another.
Ratio of sums, never mean of ratios
Over any window, the SLI is total good / total valid. Averaging daily (or hourly, or per-instance) percentages gives a quiet day the same weight as a busy one:
day valid bad daily SLI
Mon 100000 100 99.90%
Tue 300000 6000 98.00% (a sale: 3x traffic, and it broke)
Wed 100000 100 99.90%
mean of daily SLIs (99.90 + 98.00 + 99.90) / 3 = 99.27%
ratio of sums 1 - 6200 / 500000 = 98.76%
The mean says "a bit under 99.3%". Users experienced 98.76% - most of them came on Tuesday. The windows mission later in this chapter finds the same effect in real data.
Low traffic
At 20 requests a minute, one error is a 5% error rate for that minute. SLIs over short windows on low-traffic services are mostly noise. Use longer windows, give synthetic probes their own SLI (not mixed into user traffic), or merge small services into one journey. The alert-design lesson comes back to this.
In short
spec vs impl write both down; the bugs live between them
valid users only: no probes, no monitors
good decide 4xx, 429, 499 explicitly; 200-with-error needs a content check
where edge for availability; synthetic for 3am; client for the real truth
types availability, latency (fraction under a limit), correctness, freshness
windows rolling 30 (or 28) days for the SLO; calendar for reports and SLAs
combine sum the counts, then divide
What you can now do:
- decide which requests are valid and which are good, and defend it
- choose where to measure an SLI and over what window
- combine SLIs across days or routes without averaging percentages