OnCallReady

Lesson 27.15 · Observability I: Prometheus & PromQL · 31 min read

PromQL II: histograms, joins, absent, predict_linear

In plain words

Imagine a race where you do not record each runner's exact time, only how many finished under 1 minute, under 2 minutes, under 5 minutes. To guess the time of the 99th of 100 runners, you find the box they fall into, say "between 1 and 2 minutes", and guess a spot inside it. Your guess can only be as precise as that box is narrow. And if two schools ran, you must pour both schools' boxes together first; averaging each school's "99th runner" gives a wrong answer.

That is histogram_quantile(0.99, sum by (le, uri) (rate(..._bucket[5m]))): sum the buckets, keep le, then estimate. The chapter also joins series (on, group_left), catches missing data with absent(), and predicts a disk filling with predict_linear.

Latency, joins and "what if it is missing"

Rates answer "how many" and "how fast". The next questions on call are harder: "how slow is it for the unluckiest 1% of users?", "which team owns this service?", "will the disk fill up tonight?", "is that job even there?". Each needs one more PromQL tool.

What you need to know already: histograms, buckets, le, _sum and _count (27.2); rate, sum by, offset (27.8); percentiles and why they do not average (0.2); latency SLIs (0.9).

Percentiles from a histogram

histogram_quantile(φ, buckets) estimates a percentile from bucket counters. φ (phi) is the percentile as a fraction: 0.99 for p99, 0.5 for the median. The canonical p99 by endpoint (it is one of the Notion questions, learn it cold):

histogram_quantile(0.99, sum by (le, uri) (rate(http_server_requests_seconds_bucket{job="orders"}[5m])))

Read it inside out:

  1. rate(..._bucket[5m]) - each bucket is a counter; per-second rate over 5 min.
  2. sum by (le, uri) - add up all instances, methods and statuses, keeping le (the bucket boundary) and whatever you want one answer per (uri).
  3. histogram_quantile(0.99, ...) - for each uri, find the bucket where the 99th percentile falls and interpolate (estimate a value in between) inside it.
$ promtool query instant http://localhost:9090 'histogram_quantile(0.99, sum by (le, uri) (rate(http_server_requests_seconds_bucket{job="orders"}[5m])))'
{uri="/api/orders"} => 0.9119796954314702
{uri="/api/checkout"} => 1.9792857142857003
{uri="/actuator/health"} => 0.024749999999999998

The answer is in seconds: 99% of /api/orders requests finished within 0.91 s, 99% of checkouts within 1.98 s, health checks within 25 ms.

Forget le in the by and there are no buckets left to work with: the result is empty. Forget the sum entirely and you get a p99 per individual series, which is rarely what you meant:

$ promtool query instant http://localhost:9090 'histogram_quantile(0.99, rate(http_server_requests_seconds_bucket{job="orders",uri="/api/orders"}[5m]))'
{... instance="localhost:8080", method="GET", status="200" ...} => 0.1739636363636351
{... instance="localhost:8080", method="GET", status="404" ...} => 0.096
{... instance="localhost:8080", method="GET", status="500" ...} => 0.04975
... 12 lines, some NaN

NaN ("not a number") means "no observations in the window" for that series - a status that had no requests in five minutes, so 0 / 0.

How good is the estimate?

histogram_quantile assumes observations are spread evenly inside a bucket. Look at the buckets themselves:

$ promtool query instant http://localhost:9090 'sum by (le) (rate(http_server_requests_seconds_bucket{job="orders",uri="/api/checkout"}[5m]))'
{le="0.025"} => 0.15789473684210525
{le="0.05"} => 0.36842105263157887
{le="0.1"} => 0.7719298245614035
{le="0.25"} => 1.4982456140350877
{le="0.5"} => 1.8385964912280701
{le="1.0"} => 2.0035087719298246
{le="2.5"} => 2.0526315789473686
{le="5.0"} => 2.0561403508771927
{le="10.0"} => 2.0561403508771927
{le="+Inf"} => 2.0561403508771927

Each line: checkouts per second that took at most le seconds. 2.056 per second in total (+Inf). The 99th percentile is the request at position 99% of 2.056 = 2.035, which lies between the 1.0 bucket (2.0035) and the 2.5 bucket (2.0526): the answer is somewhere in 1-2.5 s, and histogram_quantile says 1.98 by drawing a straight line inside the bucket. The real value could be anywhere in that bucket. The precision of a percentile is the width of the bucket it lands in. Put boundaries where your SLO threshold is.

If the percentile falls in the +Inf bucket, you get the upper bound of the last finite bucket (here 10) - the true value is "more than 10 s" and the query cannot tell you how much more.

For an SLO you usually do not need a percentile at all; you need the fraction of requests under the threshold (the latency SLI, 0.9), which one bucket gives you exactly:

$ promtool query instant http://localhost:9090 'sum(rate(http_server_requests_seconds_bucket{job="orders",uri="/api/checkout",le="0.5"}[5m])) / sum(rate(http_server_requests_seconds_count{job="orders",uri="/api/checkout"}[5m]))'
{} => 0.8941979522184301

Requests under 0.5 s divided by all requests: 89.4% of checkouts completed in under 500 ms. No interpolation, no estimate.

Averaging percentiles is wrong

orders runs on two instances: this box, and 10.0.3.21 on a slower node.

$ promtool query instant http://localhost:9090 'histogram_quantile(0.99, sum by (le, instance) (rate(http_server_requests_seconds_bucket{job="orders",uri="/api/orders"}[5m])))'
{instance="localhost:8080"} => 0.21749681528662437
{instance="10.0.3.21:8080"} => 1.0353571428571402

$ promtool query instant http://localhost:9090 'avg(histogram_quantile(0.99, sum by (le, instance) (rate(http_server_requests_seconds_bucket{job="orders",uri="/api/orders"}[5m]))))'
{} => 0.6264269790718823

$ promtool query instant http://localhost:9090 'histogram_quantile(0.99, sum by (le) (rate(http_server_requests_seconds_bucket{job="orders",uri="/api/orders"}[5m])))'
{} => 0.9119796954314702

First, one p99 per instance: 217 ms here, 1035 ms on the slow node. Second, the average of those two: 626 ms. Third, the real p99 of the service (both instances' buckets added first): 912 ms.

The average treats both instances as equally important; the real percentile is dominated by where the slow requests are, and the slow node takes 28% of the traffic. The average is wrong in both directions depending on the traffic split, and it can hide a completely broken instance behind a healthy one. Sum the buckets first, then take the percentile. This is also why summaries (27.2) cannot be combined at all.

Average latency, the cheap way

$ promtool query instant http://localhost:9090 'sum by (uri) (rate(http_server_requests_seconds_sum{job="orders"}[5m])) / sum by (uri) (rate(http_server_requests_seconds_count{job="orders"}[5m]))'
{uri="/api/orders"} => 0.10703985209531629
{uri="/api/checkout"} => 0.2228327645051195
{uri="/actuator/health"} => 0.003195876288659786

Seconds of work per second divided by requests per second is seconds per request: 107 ms average for /api/orders. Correct, cheap, and it adds up across instances fine. Use it alongside p99, never instead: an average of 107 ms says nothing about the one in a hundred users waiting 900 ms.

Binary operators and vector matching

A binary operator is + - * / > < and friends between two things. Between a vector and a scalar (x * 100), it applies to every sample. Between two vectors, Prometheus has to match samples from the left to samples from the right. By default it matches series with exactly the same labels (ignoring the metric name):

$ promtool query instant http://localhost:9090 'rate(http_server_requests_seconds_count{job="orders",status=~"5.."}[5m]) / rate(http_server_requests_seconds_count{job="orders"}[5m])'
{..., instance="localhost:8080", method="GET", status="500", uri="/api/orders"} => 1
{..., instance="localhost:8080", method="POST", status="500", uri="/api/orders"} => NaN
...

Useless: each 5xx series matched only itself on the right (same labels, including status="500"), so the "ratio" is 1, or 0/0 = NaN. The labels that differ (status, outcome...) have to go before you divide. Aggregate both sides to the same shape:

$ promtool query instant http://localhost:9090 'sum by (instance, uri) (rate(http_server_requests_seconds_count{job="orders",status=~"5.."}[5m])) / sum by (instance, uri) (rate(http_server_requests_seconds_count{job="orders"}[5m]))'
{instance="localhost:8080", uri="/api/orders"} => 0.00028563267637817766
{instance="localhost:8080", uri="/api/checkout"} => 0
{instance="10.0.3.21:8080", uri="/api/orders"} => 0
{instance="10.0.3.21:8080", uri="/api/checkout"} => 0

Now both sides have exactly {instance, uri} and match one to one. That is the error ratio (the errors SLI, 0.8) per instance and endpoint: 0.029% on /api/orders here, zero elsewhere. It is the most important query shape in this whole block; the SLO alerts in the next chapter are this query at several window sizes.

on, ignoring, group_left

When the two sides cannot be aggregated to the same labels, tell Prometheus what to match on. on (labels) matches only on those; ignoring (labels) matches on everything except those.

group_left allows many series on the left to match one on the right, and is how you attach labels from an info metric: a series whose value is always 1 and whose labels carry facts (a version, a hostname, an owner). The node exporter publishes node_uname_info{nodename="oncall-lab", release="7.0.0-31-generic", ...} 1 - the same facts uname -a prints (1.1):

$ promtool query instant http://localhost:9090 'rate(node_network_receive_bytes_total{device="enp0s1"}[5m]) * node_uname_info'
# (nothing: the left has device="enp0s1", the right does not, no exact match)

$ promtool query instant http://localhost:9090 'rate(node_network_receive_bytes_total{device="enp0s1"}[5m]) * on (instance) group_left (nodename) node_uname_info'
{device="enp0s1", instance="localhost:9100", job="node", nodename="oncall-lab"} => 154914.15838596443

Multiplying by an info metric (value 1) changes nothing numerically; on (instance) matches the two by instance only; group_left (nodename) copies nodename from the right onto the result. The result: 155 KB/s received on enp0s1 (1.13), now labelled with the machine's name. In Kubernetes this is how you put the owning team or the node's zone onto a pod's metrics (* on (namespace, pod) group_left (owner_name) kube_pod_owner).

Get it the wrong way round and Prometheus refuses:

found duplicate series for the match group {instance="localhost:8080"} on the right hand-side of the operation: [...];many-to-many matching not allowed: matching labels must be unique on one side

The "one" side must have a single series per match key (here: per instance). group_left means the left is the "many" side; group_right the opposite.

Comparisons filter, bool compares

$ promtool query instant http://localhost:9090 'up == 0'
# (nothing: every target is up)

$ promtool query instant http://localhost:9090 'up == bool 0'
{instance="localhost:9090", job="prometheus"} => 0
{instance="localhost:9100", job="node"} => 0
...

Without bool, a comparison is a filter: series that pass keep their value, the others disappear. up == 0 returned nothing because no target is down. That is how alert conditions work - the alert fires for every series left in the result. With bool every series stays and the value becomes 1 (true) or 0 (false): here every "is it 0?" answer is 0, false.

absent(): alert on nothing

$ promtool query instant http://localhost:9090 'absent(up{job="payments"})'
{job="payments"} => 1

$ promtool query instant http://localhost:9090 'absent(up{job="node"})'
# (nothing: the series exists)

There is no payments job on this box, so the first returns 1; the node job exists, so the second returns nothing. absent returns a single series with value 1 when its argument is empty, with labels taken from the = matchers. It is the only way to notice a job disappearing entirely (from service discovery, a bad relabel rule, a removed config block) - up == 0 cannot, because there is no up left to be zero (27.2, "lookback and staleness").

predict_linear and deriv: gauges over time

$ promtool query instant http://localhost:9090 'node_filesystem_avail_bytes{mountpoint="/data"}'
{device="/dev/vdb1", fstype="ext4", instance="localhost:9100", job="node", mountpoint="/data"} => 6654279866.666667

$ promtool query instant http://localhost:9090 'predict_linear(node_filesystem_avail_bytes{mountpoint="/data"}[1h], 4 * 3600)'
{... mountpoint="/data"} => 5253043200.001381

$ promtool query instant http://localhost:9090 'node_filesystem_avail_bytes{mountpoint="/data"} / -deriv(node_filesystem_avail_bytes{mountpoint="/data"}[1h]) / 3600'
{... mountpoint="/data"} => 19.01228374603363

Three steps. The gauge: 6.65 GB free on /data right now (the same number df shows, 4.19). predict_linear(v[1h], 4*3600) draws a straight line through the last hour of samples and says where it will be in 4 x 3600 seconds = four hours: 5.25 GB free. deriv(v[1h]) is that line's slope in bytes per second (negative while the disk fills); free space divided by the negated slope, divided by 3600, is "hours until full": 19 hours.

The classic alert is predict_linear(...[6h], 4 * 3600) < 0: "will be full within four hours at the current rate", which fires before the disk is full rather than after.

Subqueries

Some questions need a value computed at many points in time: "what was the worst 5-minute error rate in the last three hours?". A subquery does that: (expr)[3h:1m] evaluates expr once a minute over the last 3 hours, producing a range vector you can feed to an _over_time function:

$ promtool query instant http://localhost:9090 'max_over_time(sum(rate(http_server_requests_seconds_count{job="orders",status=~"5.."}[5m]))[3h:1m])'
{} => 1.7403508771929825

The inner expression is the orders 5xx rate; [3h:1m] evaluates it 180 times; max_over_time keeps the highest: 1.74 errors per second, during the outage 1.5 hours ago. Handy for one-off questions; expensive (180 evaluations of the inner query), so do not put them in dashboards that refresh every 10 s. A recording rule (27.27) for the inner part plus max_over_time over the recorded series is the cheap equivalent.

The functions worth knowing by name

rate irate increase                    counters
delta idelta deriv predict_linear       gauges
avg_over_time max_over_time min_over_time sum_over_time count_over_time quantile_over_time last_over_time
histogram_quantile                      histograms
absent absent_over_time                 missing data
changes resets                          "how many times did it change / reset"
label_replace label_join                rewrite labels in a query
clamp_min clamp_max abs round ceil floor sqrt
time timestamp vector scalar
sort sort_desc topk bottomk

The *_over_time family takes a range vector of a gauge and summarises it: avg_over_time(node_load1[10m]) is the average load over ten minutes. time() is "now" as a Unix timestamp, handy for "how long ago".

label_replace deserves one example because it rescues joins where the label names differ:

label_replace(up{job="node"}, "host", "$1", "instance", "(.+):\\d+")

The arguments are: the series, the label to write (host), what to write ($1), the label to read (instance), and an anchored regex on it. It adds host="localhost" from instance="localhost:9100", so it can be matched against something labelled by host name - the same idea as a replace relabel rule (27.4), but inside a query.

What you can now do

Why it helps

"p99 by endpoint" is the query every latency panel and SLO starts from, and interviewers ask you to write it. Getting it wrong is subtle: forget le and the panel is empty; average per-instance p99s and a broken slow node hides behind a healthy one, as the orders example shows, 626 ms reported against a real 912 ms.

The rest is what separates alerts that work from alerts that do not. Division without matching labels makes every error ratio 1. absent(up{job="payments"}) is the only way to notice that a job vanished from discovery. predict_linear(...[6h], 4*3600) < 0 pages you hours before /data is full instead of after writes fail. group_left puts the owning team on a pod's metrics so alerts route correctly. You will write all of these in the next chapter's rules.

Commands in this lesson

promtool

FAQ

Why must le be in the by clause for histogram_quantile?

histogram_quantile needs the set of cumulative buckets, identified by the le label, for each group it computes. If you aggregate le away, each group has just one number and no buckets, so the result is empty. Keep le plus whatever you want one answer per: sum by (le, uri) gives one p99 per uri, sum by (le) one for the whole service.

Can I average p99 values from several instances?

No. The average of per-instance percentiles is not the percentile of the combined traffic. It ignores how much traffic each instance handles and can hide a slow instance behind a fast one, or exaggerate it. Sum the bucket rates across instances first, then run histogram_quantile once. This is also why summaries, which compute quantiles inside each app, cannot be aggregated.

Why is my division returning 1 or NaN for every series?

Binary operators between two vectors match series with exactly the same labels. If you divide 5xx requests by all requests without aggregating, each 5xx series matches only itself on the right, giving 1, or 0/0 = NaN when there was no traffic. Aggregate both sides to the same labels, sum by (instance, uri), or use on (...) or ignoring (...) to say what to match on.

When should I use absent() instead of up == 0?

up == 0 fires when Prometheus tried to scrape a target and failed. If the target disappeared from service discovery, a relabel rule dropped it, or someone deleted the job, there is no up series at all, and up == 0 returns nothing. absent(up{job="payments"}) returns 1 exactly when no such series exists, so it catches the case where monitoring itself silently lost the target.

What does group_left do?

It allows many-to-one matching: several series on the left can match one series on the right, and group_left (label) copies that label from the right onto the result. The common use is joining with an info metric of value 1, like node_uname_info or kube_pod_owner, to add labels such as nodename or owner. The right side must be unique per match key, otherwise Prometheus refuses with a many-to-many error.

In an interview Mid

How do you calculate p99 latency per endpoint in Prometheus?

histogram_quantile(0.99, sum by (le, uri) (rate(http_server_requests_seconds_bucket{job="orders"}[5m])))
  1. rate(..._bucket[5m]) - each bucket is a counter.
  2. sum by (le, uri) - add up all instances, methods and statuses, keeping le (forget it and the result is empty) and the label you want one answer per.
  3. histogram_quantile(0.99, ...) - find the bucket the 99th percentile falls in and interpolate inside it. The answer is in seconds.

What to say about it:

Also asked: How would you alert before a disk fills up? · Why is it wrong to average p99 latencies across instances? · How do you detect that a whole job has disappeared from Prometheus?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.