OnCallReady

Lesson 0.2 · SRE Fundamentals · 26 min read

Percentiles by hand, and why they do not average

In plain words

Line up everyone in your class by height, shortest to tallest. The person exactly in the middle is the median, the 50th percentile. The person 9 places from the tall end in a line of 100 is at the 90th percentile. You never add up heights and divide; you just line people up and point. If one giant visits the class, the average height jumps, but the person in the middle stays the same.

A latency percentile is exactly that: sort -n all request durations, then pick the value at rank ceil(N x q). p50 of checkout was 42 ms, p99 was 3 s, while the mean, 322 ms, described almost nobody. Percentiles from different servers or minutes can't be averaged; you merge the raw values or histogram bucket counts instead.

Why this lesson

A dashboard says "average response time: 322 ms". Everyone relaxes. Meanwhile one customer in ten is waiting more than a second, and one in a hundred is staring at a spinner for three. This lesson shows how to get the numbers that tell the truth, and why two teams can look at the same traffic and still disagree about "the p99".

What you need to know already: 0.1 The four golden signals (latency, p50/p99, the terminal basics: commands, flags, pipes).

A percentile is a position in a sorted list

The p99 of a set of durations is the value that 99% of them are at or below. There is no formula to apply to a mean or a sum: you sort every value and pick one. The simplest definition, and the one this chapter uses, is nearest rank:

rank(q) = ceil(N x q)          N = how many values; q = 0.99 for p99
p(q)    = sorted[rank(q)]      the value at that position, counting from 1

(ceil, "ceiling", rounds up to the next whole number: ceil(9.9) = 10.)

By hand, ten requests (ms): 12 15 9 30 11 14 900 13 10 16. Sorted:

rank   1   2   3   4   5   6   7   8   9   10
ms     9  10  11  12  13  14  15  16  30  900

p50  = sorted[ceil(10 x 0.5)]  = sorted[5]  = 13
p90  = sorted[ceil(10 x 0.9)]  = sorted[9]  = 30
p99  = sorted[ceil(10 x 0.99)] = sorted[10] = 900
mean = 1030 / 10 = 103

The mean (103 ms) is slower than 9 of the 10 requests and faster than the tenth. It describes nobody. The median (the middle value, p50: 13 ms) describes the typical request; the p99 (900 ms) describes the unlucky one. With 10 values p99 is just the maximum - a percentile is only meaningful when you have many more values than 1/(1-q): p99 wants hundreds of requests, p99.9 thousands.

The pipeline, on the box

Every percentile in this chapter is the same three steps: print one number per request, sort numerically, pick a rank. You can give that recipe a name with a shell function (a named, reusable mini-command you define once in the terminal):

$ pct() { sort -n | awk -v q=$1 '{a[NR]=$1} END {i=int(NR*q); if (i<NR*q) i++; print a[i]}'; }

Read it piece by piece:

Now checkout's whole log - its valid requests only (real customers; the health check is left out, see below). It is JSON (a text format where each record looks like {"status": 200, "request_time": 0.042, ...}), so we use jq (a command that reads JSON and picks fields out of it) to print one number per request:

$ for q in 0.5 0.9 0.95 0.99 0.999; do printf '%s ' $q; jq -r 'select(.uri != "/healthz") | .request_time*1000' /var/log/nginx/checkout.access.log | pct $q; done
0.5 42
0.9 1133
0.95 2452
0.99 3056
0.999 5003

$ jq -r 'select(.uri != "/healthz") | .request_time' /var/log/nginx/checkout.access.log | awk '{s+=$1} END {printf "mean %.0f ms over %d\n", 1000*s/NR, NR}'
mean 322 ms over 5400

The pieces:

Read the output line by line. Half of all requests finished in 42 ms or less. One in ten took more than 1.1 s. One in a hundred took 3 s - that is checkout's database connection pool giving up after 3000 ms of waiting for a free connection. One in a thousand hit 5 s, the time after which nginx stops waiting for checkout. The mean, 322 ms, is a number almost no request experienced: it sits in the gap between the fast majority and the incident's slow tail.

The classic mistake: sorting as text

$ printf '120\n9\n45\n1000\n30\n' | sort
1000
120
30
45
9
$ printf '120\n9\n45\n1000\n30\n' | sort -n
9
30
45
120
1000

(printf '120\n9\n...' prints the values one per line; \n means "new line".)

Plain sort compares characters, left to right: "1000" < "120" because '0' < '2'. Every percentile computed from a text sort is wrong, and it is wrong silently. Always sort -n (or sort -g, "general numeric", if values can be written like 1e-05).

The second classic mistake is a[int(NR*q)] without rounding up: for N = 5400 and q = 0.99 it is still right (5346 is a whole number), but for N = 250 it picks rank 247 instead of 248, and for q = 0.5 on an odd N it picks the value below the median. Off by one in the tail is off by a lot: the tail is where values jump.

Successes and failures are different populations

$ jq -r 'select(.uri != "/healthz" and .status < 500) | .request_time' /var/log/nginx/checkout.access.log | awk '{s+=$1} END {printf "mean ok %.0f ms over %d\n", 1000*s/NR, NR}'
mean ok 194 ms over 5184

$ jq -r 'select(.status >= 500) | .status' /var/log/nginx/checkout.access.log | sort | uniq -c
    170 500
      3 502
     43 504

The second command prints the status of every failed request, sorts them so equal codes sit next to each other, and uniq -c counts each group: 170 times 500, 3 times 502, 43 times 504.

The 500s are checkout giving up after 3 s of waiting for a database connection; the 504s are nginx giving up after 5 s; the 502s are checkout closing the connection on nginx. None of them is a fast failure. In other services it is the other way round: a service that fails fast (503 Service Unavailable in 2 ms, refusing work it knows it cannot do) makes a mixed latency number look better the more it fails. Either way, one number for both populations lies. Keep latency for successes as its own metric, and look at the latency of errors to learn how things fail (timeouts vs rejections).

Percentiles over time

A single p99 over 4.5 hours hides when. Group by minute. With 20 requests a minute, the nearest-rank p99 is rank ceil(20 x 0.99) = 20: the slowest request of the minute, so a per-minute maximum is the p99 here:

$ jq -r 'select(.uri != "/healthz") | [.time[11:16], .request_time*1000] | @tsv' /var/log/nginx/checkout.access.log | awk -F'\t' '{n[$1]++; if ($2 > mx[$1]) mx[$1]=$2} END {for (m in n) print m, n[m], mx[m]}' | sort | sed -n '176,184p'
18:27 20 110
18:28 20 164
18:29 20 139
18:30 20 126
18:31 20 147
18:32 20 3029
18:33 20 3046
18:34 20 3059
18:35 20 3045

The pieces: .time[11:16] cuts characters 11 to 15 out of a timestamp like 2026-09-22T18:32:02 - the HH:MM. [a, b] | @tsv prints the two values separated by a tab. awk -F'\t' splits columns on tabs (-F = field separator), counts requests per minute in n and keeps the largest duration in mx. for (m in n) visits the minutes in no particular order, hence the sort. sed -n '176,184p' prints only lines 176 to 184 (-n = print nothing unless told; p = print).

Your clock times will differ; the shape will not. A release (deploy: putting a new version of the program into service) ran at 18:27-18:30; at 18:32 the per-minute p99 jumps twentyfold, to the 3 s pool timeout. This is what a latency graph on a dashboard is: one percentile per time bucket.

Percentiles do not average

Busy services run as several identical copies (instances) behind nginx, which spreads requests across them. The p99 of all of them together is not the average of each instance's p99, and the p99 of the day is not the average of the hourly p99s. Suppose instance A handled 1000 requests with p99 = 50 ms, and instance B handled 10 requests with p99 = 2000 ms. The average of the two p99s is 1025 ms. The real overall p99 is rank 1000 of the 1010 merged values, which is almost certainly one of A's (about 50 ms). The average is meaningless because it ignores how many requests each percentile describes.

The same holds for any grouping: across routes, regions, minutes. What you can merge is the raw data (every duration) or counts - which is exactly why monitoring tools store latency as a histogram.

Histograms: counts in buckets

A histogram does not keep durations. It keeps, for a fixed set of limits ("buckets"), how many requests took less than or equal to each limit (cumulative: each bucket includes the ones before it), plus the total. Counts add up across instances and minutes, so histograms combine correctly. Checkout's valid requests, bucketed:

le_0.05=3082
le_0.1=4208
le_0.25=4634
le_0.5=4696
le_1=4832
le_2.5=5140
le_5=5366
le_inf=5400

le means "less than or equal to", in seconds. Read it: 3082 requests in 50 ms or less; 5140 in 2.5 s or less; all 5400 in "infinity" or less. The number between two limits is the difference: 5366 - 5140 = 226 requests took more than 2.5 s but at most 5 s.

To estimate p99 from the buckets alone: rank = 0.99 x 5400 = 5346. The first bucket whose count reaches 5346 is le_5 (5366). The previous limit is 2.5 s with 5140 below it. Assume the 226 requests in the bucket are spread evenly, and interpolate (estimate a value between two known points by assuming a straight line between them):

p99 ~ 2.5 + (5 - 2.5) x (5346 - 5140) / (5366 - 5140)
    = 2.5 + 2.5 x 206 / 226
    = 4.78 s

The exact p99 was 3.06 s. The estimate is off by more than a second because the bucket is wide (2.5 to 5 s) and the values in it are not spread evenly - most are the 3 s pool timeout. Monitoring tools do exactly this interpolation, and it is the most common reason two dashboards disagree about a p99. Two rules follow:

If the rank lands in the inf bucket, there is no upper limit to interpolate to, and the tool returns the highest finite limit - the estimate is capped by your bucket layout.

Later (Ch 27): you will build these histograms in a real monitoring tool, Prometheus, and ask it for percentiles with histogram_quantile().

The tail is the user's median

$ awk 'BEGIN {for (n=1; n<=1000; n*=10) printf "%4d backends: %.1f%% of screens hit a p99\n", n, 100*(1-0.99^n)}'
   1 backends: 1.0% of screens hit a p99
  10 backends: 9.6% of screens hit a p99
 100 backends: 63.4% of screens hit a p99
1000 backends: 100.0% of screens hit a p99

(BEGIN {...} runs before any input, so awk works here as a calculator; the for repeats with n = 1, 10, 100, 1000.)

A screen that fans out to 100 backend calls (or a checkout flow that makes 100 requests over a session) waits for the slowest one. If each has a 1% chance of being at its p99, 63% of screens include at least one. Your p99 is your users' typical experience. That is why SRE teams care about p99 and p99.9, not just the median, and why the tail is where engineering effort pays off.

In short

percentile    sort -n, pick rank ceil(N x q). Never an average of anything.
mean          sum / count. Fine for cost and capacity, useless for experience.
separate      successes vs failures, per route when routes differ
combine       merge raw values or bucket counts - never percentiles
histograms    exact for "fraction under a limit", approximate for percentiles

What you can now do:

Why it helps

You will read latency graphs every day, and two dashboards will disagree about p99. Knowing that monitoring systems estimate percentiles by interpolating inside histogram buckets explains it: a p99 estimated at 4.78 s from a 2.5 to 5 s bucket when the real value is 3.06 s. That is also why you will put bucket bounds at SLO thresholds when reviewing an app's measurements. Situations: someone averages per-server p99s into one number in a report; someone sorts durations without -n and gets silently wrong percentiles; someone reports the day's p99 as the average of hourly p99s. And in an interview, "why can't you average percentiles?" is a quick way to show you understand metrics.

Commands in this lesson

jq printf awk

FAQ

Why does a plain sort break percentiles?

Plain sort compares text character by character, so "1000" comes before "120" because '0' is less than '2'. Every percentile picked from a text-sorted list is wrong, and nothing warns you. Always use sort -n, or sort -g if values can be in scientific notation like 1e-05.

How many requests do I need for a meaningful p99?

Many more than 100. With 10 values, the p99 is simply the maximum. A percentile is only meaningful when N is large compared with 1/(1-q): p99 wants hundreds of requests at least, p99.9 thousands. Per-minute p99s on a low-traffic service are close to per-minute maximums, which is fine as long as you know that is what you are looking at.

Why do a monitoring dashboard and my own calculation give different p99s?

Nearest rank picks an actual value from the sorted data. A monitoring system that stores a histogram only has cumulative counts per bucket, so it finds the bucket containing the rank and interpolates linearly inside it, assuming values are spread evenly. With wide buckets and clustered values, like many requests at the 3-second timeout, the estimate can be off by a lot. If the rank lands in the last, unbounded bucket, it returns the highest finite bound.

If quantiles from histograms are approximate, what is exact?

Ratios against a bucket bound. The fraction of requests at or under 0.3 seconds is exactly le_0.3 / total, with no interpolation. That is why latency SLIs are written as "fraction of requests under 300 ms" rather than "p99 under X", and why the SLO threshold should be one of the bucket bounds.

Can I compute a service p99 from each server's p99?

No. The average of p99s ignores how many requests each describes: a server with 10 requests at 2 s and one with 1000 at 50 ms average to about 1 s, while the true p99 of all requests is about 50 ms. Merge raw durations, or add up histogram bucket counts across servers and compute the percentile from the sum. Counts add up correctly; percentiles never do.

In an interview Junior

How do you calculate a percentile, and why is the mean not enough?

Sort all the values as numbers and pick the one at a rank. With nearest rank, the rank is ceil(N x q), counting from 1: for ten requests of 9 10 11 12 13 14 15 16 30 900 ms, p50 is the 5th value (13 ms) and p90 the 9th (30 ms).

The mean of those ten is 103 ms: slower than nine of them and far faster than the tenth, so it describes nobody. A few slow requests drag it up, and it hides two groups of values (fast and slow) behind one "normal" number.

On a box the recipe is sort -n | awk: print one number per request, sort numerically, pick the rank, rounding up. Two classic mistakes: plain sort sorts as text (1000 before 120), and averaging p99s from several instances or hours - percentiles do not average, you need the underlying values.

Also asked: Why can't you average the p99 of five instances to get the overall p99? · What is a histogram, and how do you estimate a percentile from one? · Why should error latency be kept apart from success latency?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.