Why this lesson
A dashboard says "average response time: 322 ms". Everyone relaxes. Meanwhile one customer in ten is waiting more than a second, and one in a hundred is staring at a spinner for three. This lesson shows how to get the numbers that tell the truth, and why two teams can look at the same traffic and still disagree about "the p99".
What you need to know already: 0.1 The four golden signals (latency, p50/p99, the terminal basics: commands, flags, pipes).
A percentile is a position in a sorted list
The p99 of a set of durations is the value that 99% of them are at or below. There is no formula to apply to a mean or a sum: you sort every value and pick one. The simplest definition, and the one this chapter uses, is nearest rank:
rank(q) = ceil(N x q) N = how many values; q = 0.99 for p99
p(q) = sorted[rank(q)] the value at that position, counting from 1
(ceil, "ceiling", rounds up to the next whole number: ceil(9.9) = 10.)
By hand, ten requests (ms): 12 15 9 30 11 14 900 13 10 16. Sorted:
rank 1 2 3 4 5 6 7 8 9 10
ms 9 10 11 12 13 14 15 16 30 900
p50 = sorted[ceil(10 x 0.5)] = sorted[5] = 13
p90 = sorted[ceil(10 x 0.9)] = sorted[9] = 30
p99 = sorted[ceil(10 x 0.99)] = sorted[10] = 900
mean = 1030 / 10 = 103
The mean (103 ms) is slower than 9 of the 10 requests and faster than the tenth. It describes nobody. The median (the middle value, p50: 13 ms) describes the typical request; the p99 (900 ms) describes the unlucky one. With 10 values p99 is just the maximum - a percentile is only meaningful when you have many more values than 1/(1-q): p99 wants hundreds of requests, p99.9 thousands.
The pipeline, on the box
Every percentile in this chapter is the same three steps: print one number per request, sort numerically, pick a rank. You can give that recipe a name with a shell function (a named, reusable mini-command you define once in the terminal):
$ pct() { sort -n | awk -v q=$1 '{a[NR]=$1} END {i=int(NR*q); if (i<NR*q) i++; print a[i]}'; }
Read it piece by piece:
pct() { ...; }defines a function calledpct. After this line, typingpct 0.99runs what is between the braces, and$1inside means "the first argument you gave it" (0.99).sort -nsorts its input as numbers.awk -v q=$1starts awk with a variable (a named value)qset to that argument (-v= "set this variable").{a[NR]=$1}runs for every line: store the value in a listaat position NR (the line number, starting at 1).END {...}runs once after the last line:int(NR*q)cuts off the decimals, theifadds one when there were decimals (that is the ceil), and it prints the value at that rank.
Now checkout's whole log - its valid requests only (real customers; the health check is left out, see below). It is JSON (a text format where each record looks like {"status": 200, "request_time": 0.042, ...}), so we use jq (a command that reads JSON and picks fields out of it) to print one number per request:
$ for q in 0.5 0.9 0.95 0.99 0.999; do printf '%s ' $q; jq -r 'select(.uri != "/healthz") | .request_time*1000' /var/log/nginx/checkout.access.log | pct $q; done
0.5 42
0.9 1133
0.95 2452
0.99 3056
0.999 5003
$ jq -r 'select(.uri != "/healthz") | .request_time' /var/log/nginx/checkout.access.log | awk '{s+=$1} END {printf "mean %.0f ms over %d\n", 1000*s/NR, NR}'
mean 322 ms over 5400
The pieces:
for q in A B C; do ...; donerepeats the commands once for each value, with$qset to it.jq -r '...':-r= print raw values (no quotes around them).select(.uri != "/healthz")keeps only the records whoseurifield is not/healthz;| .request_time*1000then prints that record's duration in ms.- Why drop
/healthz? It is a health check (an address that a monitoring tool calls every few seconds just to ask "are you alive?"). It answers in 2 ms and is not a customer, so it would pull every number down.
Read the output line by line. Half of all requests finished in 42 ms or less. One in ten took more than 1.1 s. One in a hundred took 3 s - that is checkout's database connection pool giving up after 3000 ms of waiting for a free connection. One in a thousand hit 5 s, the time after which nginx stops waiting for checkout. The mean, 322 ms, is a number almost no request experienced: it sits in the gap between the fast majority and the incident's slow tail.
The classic mistake: sorting as text
$ printf '120\n9\n45\n1000\n30\n' | sort
1000
120
30
45
9
$ printf '120\n9\n45\n1000\n30\n' | sort -n
9
30
45
120
1000
(printf '120\n9\n...' prints the values one per line; \n means "new line".)
Plain sort compares characters, left to right: "1000" < "120" because '0' < '2'. Every percentile computed from a text sort is wrong, and it is wrong silently. Always sort -n (or sort -g, "general numeric", if values can be written like 1e-05).
The second classic mistake is a[int(NR*q)] without rounding up: for N = 5400 and q = 0.99 it is still right (5346 is a whole number), but for N = 250 it picks rank 247 instead of 248, and for q = 0.5 on an odd N it picks the value below the median. Off by one in the tail is off by a lot: the tail is where values jump.
Successes and failures are different populations
$ jq -r 'select(.uri != "/healthz" and .status < 500) | .request_time' /var/log/nginx/checkout.access.log | awk '{s+=$1} END {printf "mean ok %.0f ms over %d\n", 1000*s/NR, NR}'
mean ok 194 ms over 5184
$ jq -r 'select(.status >= 500) | .status' /var/log/nginx/checkout.access.log | sort | uniq -c
170 500
3 502
43 504
The second command prints the status of every failed request, sorts them so equal codes sit next to each other, and uniq -c counts each group: 170 times 500, 3 times 502, 43 times 504.
The 500s are checkout giving up after 3 s of waiting for a database connection; the 504s are nginx giving up after 5 s; the 502s are checkout closing the connection on nginx. None of them is a fast failure. In other services it is the other way round: a service that fails fast (503 Service Unavailable in 2 ms, refusing work it knows it cannot do) makes a mixed latency number look better the more it fails. Either way, one number for both populations lies. Keep latency for successes as its own metric, and look at the latency of errors to learn how things fail (timeouts vs rejections).
Percentiles over time
A single p99 over 4.5 hours hides when. Group by minute. With 20 requests a minute, the nearest-rank p99 is rank ceil(20 x 0.99) = 20: the slowest request of the minute, so a per-minute maximum is the p99 here:
$ jq -r 'select(.uri != "/healthz") | [.time[11:16], .request_time*1000] | @tsv' /var/log/nginx/checkout.access.log | awk -F'\t' '{n[$1]++; if ($2 > mx[$1]) mx[$1]=$2} END {for (m in n) print m, n[m], mx[m]}' | sort | sed -n '176,184p'
18:27 20 110
18:28 20 164
18:29 20 139
18:30 20 126
18:31 20 147
18:32 20 3029
18:33 20 3046
18:34 20 3059
18:35 20 3045
The pieces: .time[11:16] cuts characters 11 to 15 out of a timestamp like 2026-09-22T18:32:02 - the HH:MM. [a, b] | @tsv prints the two values separated by a tab. awk -F'\t' splits columns on tabs (-F = field separator), counts requests per minute in n and keeps the largest duration in mx. for (m in n) visits the minutes in no particular order, hence the sort. sed -n '176,184p' prints only lines 176 to 184 (-n = print nothing unless told; p = print).
Your clock times will differ; the shape will not. A release (deploy: putting a new version of the program into service) ran at 18:27-18:30; at 18:32 the per-minute p99 jumps twentyfold, to the 3 s pool timeout. This is what a latency graph on a dashboard is: one percentile per time bucket.
Percentiles do not average
Busy services run as several identical copies (instances) behind nginx, which spreads requests across them. The p99 of all of them together is not the average of each instance's p99, and the p99 of the day is not the average of the hourly p99s. Suppose instance A handled 1000 requests with p99 = 50 ms, and instance B handled 10 requests with p99 = 2000 ms. The average of the two p99s is 1025 ms. The real overall p99 is rank 1000 of the 1010 merged values, which is almost certainly one of A's (about 50 ms). The average is meaningless because it ignores how many requests each percentile describes.
The same holds for any grouping: across routes, regions, minutes. What you can merge is the raw data (every duration) or counts - which is exactly why monitoring tools store latency as a histogram.
Histograms: counts in buckets
A histogram does not keep durations. It keeps, for a fixed set of limits ("buckets"), how many requests took less than or equal to each limit (cumulative: each bucket includes the ones before it), plus the total. Counts add up across instances and minutes, so histograms combine correctly. Checkout's valid requests, bucketed:
le_0.05=3082
le_0.1=4208
le_0.25=4634
le_0.5=4696
le_1=4832
le_2.5=5140
le_5=5366
le_inf=5400
le means "less than or equal to", in seconds. Read it: 3082 requests in 50 ms or less; 5140 in 2.5 s or less; all 5400 in "infinity" or less. The number between two limits is the difference: 5366 - 5140 = 226 requests took more than 2.5 s but at most 5 s.
To estimate p99 from the buckets alone: rank = 0.99 x 5400 = 5346. The first bucket whose count reaches 5346 is le_5 (5366). The previous limit is 2.5 s with 5140 below it. Assume the 226 requests in the bucket are spread evenly, and interpolate (estimate a value between two known points by assuming a straight line between them):
p99 ~ 2.5 + (5 - 2.5) x (5346 - 5140) / (5366 - 5140)
= 2.5 + 2.5 x 206 / 226
= 4.78 s
The exact p99 was 3.06 s. The estimate is off by more than a second because the bucket is wide (2.5 to 5 s) and the values in it are not spread evenly - most are the 3 s pool timeout. Monitoring tools do exactly this interpolation, and it is the most common reason two dashboards disagree about a p99. Two rules follow:
- put bucket limits at your targets (a "300 ms" latency target needs a 0.3 bucket, otherwise every answer about it is an estimate)
- percentiles from wide buckets are guesses; ratios against a limit ("% of requests <= 0.3 s") are exact, which is why latency targets are written as "fraction under a threshold", not "p99 under X"
If the rank lands in the inf bucket, there is no upper limit to interpolate to, and the tool returns the highest finite limit - the estimate is capped by your bucket layout.
Later (Ch 27): you will build these histograms in a real monitoring tool, Prometheus, and ask it for percentiles with
histogram_quantile().
The tail is the user's median
$ awk 'BEGIN {for (n=1; n<=1000; n*=10) printf "%4d backends: %.1f%% of screens hit a p99\n", n, 100*(1-0.99^n)}'
1 backends: 1.0% of screens hit a p99
10 backends: 9.6% of screens hit a p99
100 backends: 63.4% of screens hit a p99
1000 backends: 100.0% of screens hit a p99
(BEGIN {...} runs before any input, so awk works here as a calculator; the for repeats with n = 1, 10, 100, 1000.)
A screen that fans out to 100 backend calls (or a checkout flow that makes 100 requests over a session) waits for the slowest one. If each has a 1% chance of being at its p99, 63% of screens include at least one. Your p99 is your users' typical experience. That is why SRE teams care about p99 and p99.9, not just the median, and why the tail is where engineering effort pays off.
In short
percentile sort -n, pick rank ceil(N x q). Never an average of anything.
mean sum / count. Fine for cost and capacity, useless for experience.
separate successes vs failures, per route when routes differ
combine merge raw values or bucket counts - never percentiles
histograms exact for "fraction under a limit", approximate for percentiles
What you can now do:
- compute a nearest-rank percentile by hand and with
sort -n | awk - explain why averaging p99s across instances or hours is wrong
- read a histogram and estimate a percentile from it