OnCallReady

Lesson 0.9 · SRE Fundamentals · 18 min read

Choosing what to count: valid events, measurement points, windows

In plain words

Imagine counting how many children in your class finished a puzzle. First you decide who counts: not the teacher, not the visiting parent who just watched. Then you decide what "finished" means: is a puzzle with one piece missing finished? Then you decide where you stand to count: at the door, you catch everyone; at the puzzle table, you miss the children who never got there. Change any of these and your "90% finished" might become 97% or 85%, without any child doing anything differently.

An SLI is exactly those decisions plus a division: which events are valid (no probes), which are good (4xx usually good, 429 and 499 need a decision, a 200 with an error body is bad), where you measure (the edge), over what window (rolling 30 days), and always a ratio of sums.

Why this lesson

Two teams measure the same service over the same day. One reports 96.7%, the other 93.7%. Both used "the access log". The difference is three small choices nobody wrote down - and on a 99.9% target, the whole allowance for failure is 0.1 points. This lesson is about making those choices on purpose.

What you need to know already: 0.8 SLIs, SLOs and SLAs (good / valid, probes), 0.2 Percentiles (jq, select, sort | uniq -c).

An SLI is two decisions and a division

SLI = good events / valid events        over a window

The SRE workbook separates the SLI specification (what you want to measure, in words: "the proportion of checkout requests that succeed") from the SLI implementation (how you actually count it: "non-5xx responses logged by the edge nginx, excluding /healthz"). Most SLI bugs live in the gap between the two. You will investigate one later in this chapter: the specification meant "searches that succeed", the implementation counted probes and treated a user giving up as a success.

Decision 1: which events are valid

Start from every request the edge saw and remove what is not a user trying to do something:

On checkout, the status classes of the valid requests:

$ jq -r 'select(.uri != "/healthz") | .status' /var/log/nginx/checkout.access.log | awk '{c[int($1/100)"xx"]++} END {for (k in c) print k, c[k]}' | sort
2xx 5057
4xx 127
5xx 216

$ jq -r 'select(.uri != "/healthz") | .status' /var/log/nginx/checkout.access.log | sort | uniq -c
   4283 200
    774 201
     40 400
     37 401
     21 404
     29 409
    170 500
      3 502
     43 504

int($1/100)"xx" turns 404 into "4xx": divide by 100, drop the decimals, and glue "xx" on (awk joins values written next to each other). c[...]++ counts one per line under that key; the END loop prints every key and its count.

Decision 2: which events are good

StatusUsuallyWhy
2xx, 3xxgoodthe service did its job
400, 401, 403, 404, 409, 422goodthe service answered correctly; the client asked for something wrong
429dependsgood if the client went over a documented limit; bad if you turned away real users because you were overloaded
499 (nginx)dependsthe client hung up first. Bad when your slowness caused it (a phone app's timeout); noise when a search box cancels the previous keystroke's request
5xxbadthe service failed
200 with an error bodybadthe classic hidden failure - only a correctness SLI or a content check sees it

The same log, three honest-looking definitions:

                                              good    valid   SLI
non-5xx over valid (the SLO's definition)     5184    5400    96.00%
non-5xx over every line (probe counted)       6264    6480    96.67%
2xx only over valid (4xx counted as bad)      5057    5400    93.65%

A 3-point spread from definitions alone - on a 99.9% SLO the whole allowance for failure is 0.1 points. That is why the definition is written down, agreed, and reviewed as carefully as code.

Decision 3: where it is measured

client (the app or browser)  closest to the user; sees the whole trip over the internet.
                             Noisy, delayed, and you do not control it.
edge / load balancer         sees every request that reached you, including those no
                             copy of the service answered (502/504 from the proxy).
                             The usual choice.
application                  sees only requests that reached the service. Misses every
                             request that died in front of it - a crashed service
                             reports nothing at all.
synthetic prober             a small program that runs the journey every minute. Works at 3am
                             with no traffic; measures one path, not your users.

(Measuring in the user's browser is called RUM, real user monitoring.) For checkout the edge nginx is right: it logged the 504s that checkout itself never knew about - checkout did not know nginx had given up on it.

SLI types, with the command that counts them

Availability - you have done it all chapter. Latency - the fraction served fast enough:

$ jq -r 'select(.uri != "/healthz") | .request_time' /var/log/nginx/checkout.access.log | awk '$1 <= 0.3 {f++} END {printf "%d of %d = %.2f%% under 300 ms\n", f, NR, 100*f/NR}'
4651 of 5400 = 86.13% under 300 ms

($1 <= 0.3 {f++}: for every line whose value is at most 0.3 seconds, add one to f.) A latency SLI keeps all valid requests in the bottom of the fraction - a 5 s 504 is a slow request as well as a failed one. Some teams use two thresholds: "90% under 100 ms and 99% under 400 ms" captures both the typical and the tail experience.

Correctness (quality) - was the answer right? The log cannot tell you: a 200 with the wrong price looks perfect. You measure it with a prober that checks the response body, or a reconciliation job ("orders whose charged amount equals the cart total / orders checked"). Freshness for data pipelines ("reads of data newer than 10 minutes / reads"), coverage ("records processed / records received"), durability for storage.

Per route, or per journey

$ jq -r 'select(.uri != "/healthz") | "\(.uri) \(.status)"' /var/log/nginx/checkout.access.log | sed -E 's,/[0-9]+ , ,' | awk '{n[$1]++; if ($2 >= 500) b[$1]++} END {for (k in n) printf "%-26s %5d %4d %7.2f%%\n", k, n[k], b[k], 100*(1-b[k]/n[k])}' | sort
/api/cart                   1870   82   95.61%
/api/cart/items             1100   43   96.09%
/api/checkout                817   31   96.21%
/api/orders                  815   33   95.95%
/api/payments/authorize      527   19   96.39%
/api/shipping/quote          271    8   97.05%

Columns: route, requests, failures, SLI. Two details. sed -E 's,/[0-9]+ , ,' (sed edits text as it flows past; s,A,B, replaces A with B, and -E enables patterns like [0-9]+ = "one or more digits") turns /api/orders/40117 into /api/orders: without it every order number is its own "route", and you get hundreds of rows with one request each. That is the cardinality problem (too many distinct values of something you group by) - it also breaks monitoring tools, so never group by an id. And "\(.uri) \(.status)" is jq's way of building a text from fields.

Here the incident hit every route equally, so one service-level SLI is enough. When routes differ (a cheap GET /cart vs an expensive POST /checkout), a single SLI lets the busy cheap route hide failures on the important one. Group routes into user journeys ("browse", "pay") and give the important ones their own SLO.

Windows

Rolling (the last 30 days, recomputed continuously) is what users experience and what the SLO itself should use: every bad day counts for exactly 30 days. Calendar (this month) resets on the 1st: an outage on the 30th is forgotten two days later. Calendar windows are for reporting and contracts (SLAs are billed monthly).

28 days is a popular alternative to 30: every window then holds the same number of each weekday, so weekly traffic patterns do not skew one window against another.

Ratio of sums, never mean of ratios

Over any window, the SLI is total good / total valid. Averaging daily (or hourly, or per-instance) percentages gives a quiet day the same weight as a busy one:

day        valid     bad     daily SLI
Mon       100000     100      99.90%
Tue       300000    6000      98.00%      (a sale: 3x traffic, and it broke)
Wed       100000     100      99.90%

mean of daily SLIs     (99.90 + 98.00 + 99.90) / 3   = 99.27%
ratio of sums          1 - 6200 / 500000              = 98.76%

The mean says "a bit under 99.3%". Users experienced 98.76% - most of them came on Tuesday. The windows mission later in this chapter finds the same effect in real data.

Low traffic

At 20 requests a minute, one error is a 5% error rate for that minute. SLIs over short windows on low-traffic services are mostly noise. Use longer windows, give synthetic probes their own SLI (not mixed into user traffic), or merge small services into one journey. The alert-design lesson comes back to this.

In short

spec vs impl    write both down; the bugs live between them
valid           users only: no probes, no monitors
good            decide 4xx, 429, 499 explicitly; 200-with-error needs a content check
where           edge for availability; synthetic for 3am; client for the real truth
types           availability, latency (fraction under a limit), correctness, freshness
windows         rolling 30 (or 28) days for the SLO; calendar for reports and SLAs
combine         sum the counts, then divide

What you can now do:

Why it helps

SLI definitions are where honest-looking numbers go wrong. On the same checkout log, three reasonable-sounding definitions gave 96.00%, 96.67% and 93.65%, a 3-point spread when the whole 99.9% budget is 0.1 points. Situations: a team's SLO looks green because probes are counted; a weekly report averages daily SLIs and says 99.27% when users experienced 98.76%; route-level SLIs explode into thousands of rows because order IDs are in the path, the same problem that overloads monitoring systems. When you review someone's SLO document, these are the questions you ask, and an interviewer probing SLO depth will ask them too.

Commands in this lesson

jq

FAQ

What is the difference between an SLI specification and an implementation?

The specification is what you want to measure, in words: "the proportion of checkout requests that succeed". The implementation is how you actually count it: "non-5xx responses logged by the edge nginx, excluding /healthz". The SRE workbook separates them because most SLI bugs live in the gap, such as an implementation counting probes, or treating a user giving up as a success.

Why must I use a ratio of sums instead of averaging daily SLIs?

Averaging percentages gives a quiet day the same weight as a busy one. With Monday and Wednesday at 99.9% on 100,000 requests and a sale day at 98% on 300,000, the mean of daily SLIs is 99.27%, but users experienced 98.76%, because most of them came on the bad day. Always sum good and valid events over the window, then divide.

Rolling or calendar window for an SLO?

Rolling, for the SLO itself: the last 30 (or 28) days, recomputed continuously, so every bad day counts for exactly 30 days, matching user experience. Calendar windows reset on the 1st, forgetting an outage on the 30th two days later; they suit reports and contracts, since SLAs are billed monthly. 28 days keeps the same number of each weekday in every window.

How do I handle an SLI per route without thousands of rows?

Normalise the path before grouping: cut the IDs out, so /api/orders/40117 becomes /api/orders (the sed text tool can do it with one substitution). Otherwise every ID is its own route, and you get one row per order - the same problem that overloads monitoring systems when every ID becomes its own series. Then group routes into user journeys and give the important ones, like "pay", their own SLO.

What do I do about low-traffic services?

At 20 requests a minute, one error is 5% for that minute, so short-window SLIs are mostly noise. Use longer windows, merge small services into one journey SLO, and give synthetic probes their own separate SLI rather than mixing them into user traffic. The alerting lessons add minimum error counts before paging.

In an interview Junior

Should health checks be included in an availability SLI?

No. An SLI counts valid events - users trying to do something - and health checks are not users. They hit a cheap address like /healthz that almost never fails, so counting them makes the service look better than it is: on checkout's log, adding the probe raised availability from 96.00% to 96.67%. They also pad the denominator, so a real outage looks smaller.

So the SLI filters them out, for example with jq 'select(.uri != "/healthz")', and the filter is written into the SLI's definition so people can review it. The same goes for other traffic that is not a user. When you combine days or routes, use the ratio of sums (all good / all valid), never the mean of the daily percentages.

Also asked: Where would you measure an SLI: at the proxy in front, or inside the service, and why? · What is the difference between a rolling and a calendar window? · Why is "ratio of sums" right and "mean of ratios" wrong when combining SLIs?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.