OnCallReady

Lesson 27.8 · Observability I: Prometheus & PromQL · 37 min read

PromQL I: selectors, rate, increase, aggregation

In plain words

Imagine a water meter on your house. The number on it only goes up: litres used since it was installed. Reading "84,906 litres" tells you nothing useful. What you want is "how fast is water flowing right now", which you get by reading the meter twice, a few minutes apart, and dividing. If the meter was replaced and restarted at zero in between, you add the old reading back on so you don't get a negative.

That is rate(). A counter like http_server_requests_seconds_count is the meter; rate(x[5m]) is litres per second over the last five minutes, reset-safe. increase(x[1h]) is litres used in the hour. Selectors like {status=~"5.."} pick which meters, and sum by (uri) adds meters together, always after the rate, never before.

Asking the questions

Prometheus now holds thousands of series. "How many requests per second is checkout serving?" is not stored anywhere as such: the data is a pile of ever-growing counters, one per instance, method and status. PromQL is the query language that turns that pile into an answer - and most wrong graphs and silent alerts come from one of the half-dozen mistakes this lesson walks through.

What you need to know already: series, labels, samples, counters and gauges (27.1, 27.2); promtool query instant (27.2); regular expressions (7.3); the golden signals traffic and errors (0.1); what "requests per second" and "error ratio" mean (0.10).

The shape of a PromQL query

A query is built from pieces you nest like function calls in any language. Here is the one you will write most, taken apart:

sum by (uri) (rate(http_server_requests_seconds_count{job="orders"}[5m]))

Read it inside out: select, take a window, compute the rate, add up. The rest of this lesson is each piece in turn.

Two kinds of vector

Every PromQL expression evaluates to one of four types. Two matter. PromQL calls a set of series a vector.

Instant vector - one sample per series, the value "now" (at the evaluation time).

$ promtool query instant http://localhost:9090 'node_cpu_seconds_total{cpu="0",mode="idle"}'
node_cpu_seconds_total{cpu="0", instance="localhost:9100", job="node", mode="idle"} => 7474.41 @[1790107223.4]

One series, one value: 7474.41 idle seconds on CPU 0. @[1790107223.4] is the evaluation time as a Unix timestamp (seconds since 1 January 1970, UTC; date -u -d @1790107223 turns it into a date).

Range vector - every sample per series inside a window ending now. You write it with [duration] after a selector (s, m, h, d, w for seconds to weeks):

$ promtool query instant http://localhost:9090 'node_cpu_seconds_total{cpu="0",mode="idle"}[1m]'
node_cpu_seconds_total{cpu="0", instance="localhost:9100", job="node", mode="idle"} =>
7433.49 @[1790107168.48]
7447.18 @[1790107183.48]
7460.67 @[1790107198.48]
7474.41 @[1790107213.48]

Four samples in a minute: the node job scrapes every 15 s. Note the sample timestamps are the scrape times, 15 s apart, not a regular clock grid.

You cannot graph a range vector and you cannot do arithmetic on one. You feed it to a function (rate, increase, avg_over_time...) that turns it back into an instant vector. The other two types are scalar (a bare number, like 3600) and string (only used as function arguments).

Selectors and matchers

http_server_requests_seconds_count{job="orders", status="500"}        =  exact
http_server_requests_seconds_count{job="orders", status!="200"}       != not equal
http_server_requests_seconds_count{job="orders", status=~"5.."}       =~ regex
http_server_requests_seconds_count{job="orders", uri!~"/actuator/.*"} !~ negative regex
{__name__=~"node_memory_.*_bytes", instance="localhost:9100"}         the name is a label too

Several matchers in one {} must all be true (AND). The last line shows that the metric name is really a label called __name__, so you can select several metrics by a pattern.

Regexes are RE2 (Go's flavour, close to ERE, 7.3) and fully anchored: status=~"5" matches only the string 5, not 500. Write 5.. (a 5 and any two characters). And status=~"5xx" matches the literal text "5xx", which no series has, so it returns nothing - silently. That exact typo turns up again in the next chapter, inside an alert.

A missing label matches the empty string, so {env=""} selects series that have no env label. A selector must contain at least one matcher that does not match everything:

$ promtool query instant http://localhost:9090 '{job=~".*"}'
query error: bad_data: invalid parameter "query": 1:1: parse error: vector selector must contain at least one non-empty matcher

.* matches the empty string too, so this would select every series in the database. Prometheus refuses.

rate(): the function you will use most

A counter's raw value is meaningless on a graph ("84906 requests since the JVM started"). What you want is how fast it grows:

$ promtool query instant http://localhost:9090 'rate(node_cpu_seconds_total{mode="idle"}[5m])'
{cpu="0", instance="localhost:9100", job="node", mode="idle"} => 0.9086666666666675
{cpu="1", instance="localhost:9100", job="node", mode="idle"} => 0.9077543859649123

rate(x[5m]) is the per-second average increase over the last five minutes. For CPU seconds that is "seconds of idle per second" = 0.91, so each core is 91% idle. Note the metric name is gone from the result: rate returns a different thing than the input (a speed, not a count), so Prometheus drops __name__.

How it is computed matters:

  1. take the first and last sample in the window;
  2. if any sample is lower than the one before it, the counter reset (the process restarted, 27.2): add the value before the drop back on, so the reset does not look like a huge negative;
  3. extrapolate - stretch the result to the edges of the window, because the first and last samples are never exactly at its edges;
  4. divide by the window length in seconds.

Step 3 is why increase() of a whole-number counter often returns something like 4871.297071129707: it is an estimate, not a count.

The window must hold at least two samples

$ promtool query instant http://localhost:9090 'rate(http_server_requests_seconds_count{instance="localhost:8080",uri="/api/checkout",status="200"}[30s])'
{..., uri="/api/checkout"} => 1.4

$ promtool query instant http://localhost:9090 'rate(http_server_requests_seconds_count{instance="localhost:8080",uri="/api/checkout",status="200"}[15s])'
# (nothing)

The first: 1.4 successful checkouts per second on this instance. The second prints nothing at all. With a 15 s scrape interval a 15 s window holds one sample, and a rate needs two. No error, just no data - the graph has gaps, or an alert built on it never fires. [30s] works but is fragile: one missed scrape (a slow target, a Prometheus restart) and it is empty again. The rule: the range is at least four times the scrape interval. With 15 s scrapes, [1m] is the minimum you should write and [5m] the usual choice.

irate(): only the last two samples

$ promtool query instant http://localhost:9090 'irate(http_server_requests_seconds_count{instance="localhost:8080",uri="/api/checkout",status="200"}[5m])'
{..., uri="/api/checkout"} => 1.4

irate ("instant rate") ignores everything in the window except the last two samples. It reacts instantly and jumps around; good for a zoomed-in live graph of a jumpy counter, bad for alerts (one noisy pair of scrapes sets it off) and bad for long time ranges on graphs (it only ever looks at two points, so a graph with one point per hour hides everything in between).

increase(): rate times the window

$ promtool query instant http://localhost:9090 'increase(http_server_requests_seconds_count{instance="localhost:8080",uri="/api/checkout",status="200"}[1h])'
{..., uri="/api/checkout"} => 4871.297071129707

About 4871 successful checkouts in the last hour. increase(x[1h]) is exactly rate(x[1h]) * 3600. Use increase when the question is a count over a period ("how many checkouts failed in the last hour", "requests per day"), rate when it is a speed ("requests per second"), and never mix a rate in one graph and an increase with a different window in another without saying so.

Why a counter reset does not break rate()

The orders app restarted twice in the last six hours (the boot and a deploy). resets(x[6h]) counts how many times a counter dropped in the window:

$ promtool query instant http://localhost:9090 'resets(http_server_requests_seconds_count{instance="localhost:8080",uri="/api/checkout",status="200"}[6h])'
{...} => 2

The naive "value now minus value four hours ago" is nonsense across a reset. offset 4h after a selector means "as it was 4 hours ago":

$ promtool query instant http://localhost:9090 'http_server_requests_seconds_count{instance="localhost:8080",uri="/api/checkout",status="200"} - http_server_requests_seconds_count{instance="localhost:8080",uri="/api/checkout",status="200"} offset 4h'
{...} => -391608

A negative number of requests: the counter was high four hours ago, went back to 0 at the restart, and has not climbed that high since. increase over the same window gives the real answer because it detects the drop and adds the pre-reset value back:

$ promtool query instant http://localhost:9090 'increase(http_server_requests_seconds_count{instance="localhost:8080",uri="/api/checkout",status="200"}[4h])'
{...} => 20408.25860271116

The one thing rate cannot see is a restart between two scrapes where the new process has already counted past the old value. Rare, and the error is small.

rate first, then sum. Never the other way

Aggregation operators (sum, avg...) combine many series into fewer. rate must see each counter on its own so it can detect each one's resets. So:

sum(rate(http_server_requests_seconds_count{job="orders"}[5m]))     right
rate(sum(http_server_requests_seconds_count{job="orders"})[5m])     not even valid
$ promtool query instant http://localhost:9090 'rate(sum(http_server_requests_seconds_count{job="orders"})[5m])'
query error: bad_data: invalid parameter "query": 1:59: parse error: ranges only allowed for vector selectors

[5m] may only follow a plain selector, and sum(...) is not one.

(You can force it with a subquery, rate(sum(x)[5m:]) - 27.15 - and then one instance restarting makes the sum drop and rate reports a reset of the whole service. Don't.)

The same rule applied to functions: rate wants a range vector, so passing an instant vector is a type error Prometheus catches for you:

$ promtool query instant http://localhost:9090 'rate(node_load1)'
query error: bad_data: invalid parameter "query": 1:1: parse error: expected type range vector in call to function "rate", got instant vector

And rate on a gauge is a logic error nobody catches: rate(node_load1[5m]) returns numbers, and they mean nothing (a gauge going down is not a reset). For gauges use deriv, delta or the *_over_time functions (27.15).

Aggregation: by and without

$ promtool query instant http://localhost:9090 'sum by (uri) (rate(http_server_requests_seconds_count{job="orders"}[5m]))'
{uri="/api/orders"} => 17.08070175438596
{uri="/api/checkout"} => 2.0561403508771927
{uri="/actuator/health"} => 0.3403508771929824

The orders service serves 17 requests per second on /api/orders, 2 on checkout and 0.34 health checks - all instances, methods and statuses added together. by (labels) keeps only those labels; everything else is summed away. without (labels) removes those and keeps the rest:

$ promtool query instant http://localhost:9090 'sum without (instance) (rate(http_server_requests_seconds_count{job="orders",uri="/api/orders",method="GET"}[5m]))'
{error="none", exception="none", job="orders", method="GET", outcome="SUCCESS", status="200", team="orders", uri="/api/orders"} => 13.617543859649121
{error="none", exception="none", job="orders", method="GET", outcome="CLIENT_ERROR", status="404", team="orders", uri="/api/orders"} => 0.04210526315789473
{error="CannotGetJdbcConnectionException", ..., status="500", ...} => 0.003508771929824561

Both instances added together, every other label kept: 13.6 successful GETs per second, a trickle of 404s, and a very small rate of 500s caused by a database connection error.

Use without in saved queries that should keep working when someone adds a label; use by when you want exactly a given shape (one series per uri). The clause can go before or after the parentheses: sum by (uri) (x) and sum(x) by (uri) are the same.

The other aggregators:

avg, min, max, count, group                  the obvious ones; group returns 1
stddev, stdvar                               spread
topk(3, x), bottomk(3, x)                    keep the 3 biggest series, labels intact
quantile(0.9, x)                             the 90th percentile ACROSS series
count_values("version", build_info)          how many series have each value

count counts series, not requests. How many targets does each job have?

$ promtool query instant http://localhost:9090 'count by (job) (up)'
{job="prometheus"} => 1
{job="node"} => 1
{job="orders"} => 2
{job="lab"} => 1

Arithmetic on the results

$ promtool query instant http://localhost:9090 '100 * (1 - avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])))'
{instance="localhost:9100"} => 9.178947368421008

CPU busy percent: 9.2%. Each piece: mode="idle" because "busy" is everything else; rate because the metric is a counter of seconds; avg by (instance) because there is one series per core and you want one per machine; 1 - ... turns idle into busy; 100 * makes it a percent. This is the query behind every "CPU %" graph you will ever see.

offset: look at the past

$ promtool query instant http://localhost:9090 'sum(rate(http_server_requests_seconds_count{job="orders"}[5m])) / sum(rate(http_server_requests_seconds_count{job="orders"}[5m] offset 1h))'
{} => 0.9574944071588365

"Traffic now divided by traffic an hour ago": 0.957, so 4% lower. offset goes right after the selector (inside the rate) and shifts it back in time. {} means the result has no labels left - sum without by removed them all. offset 1d or offset 1w is the classic "same time yesterday / last week" comparison. On this box those return nothing, because this Prometheus only has six hours of data - a reminder that offset beyond your retention is silent.

Reading an error message

$ promtool query instant http://localhost:9090 'sum(rate(foo[5m]'
query error: bad_data: invalid parameter "query": 1:17: parse error: unexpected end of input in function call, expected ")"

1:17 is line 1, column 17: where the parser gave up (a closing parenthesis is missing). bad_data means the query did not parse (HTTP 400); execution errors (HTTP 422) are queries that parse but cannot run, like a join that matches too many series (27.15).

What you can now do

Why it helps

Every alert and panel you will write starts with this lesson's shapes. At 3am you need "which endpoint is failing?" as sum by (uri) (rate(...{status=~"5.."}[5m])) without thinking. When a panel has gaps or an alert never fires, you check whether the range is at least four times the scrape interval. When someone's query uses status=~"5xx" and shows zero errors during an outage, you spot the anchored-regex bug, the incident from the next chapter.

Rate-then-sum is a classic interview question and a common review finding: summing counters first means one pod restart looks like a service-wide reset. Knowing that increase() returns an extrapolated estimate like 4871.29 saves an argument with a manager who expects an exact count. The CPU busy query is one you should be able to derive from scratch.

Commands in this lesson

promtool

FAQ

What is the difference between rate, irate and increase?

rate(x[5m]) is the per-second average increase over the whole window, handling resets and extrapolating to the window edges; use it for graphs and alerts. irate(x[5m]) uses only the last two samples, so it is jumpy, fine for zoomed-in live graphs, bad for alerts. increase(x[1h]) is rate(x[1h]) * 3600, the estimated total increase over the window.

Why does rate with a short range return nothing?

A rate needs at least two samples in the window. With a 15 second scrape interval, [15s] usually holds one sample, so the result is empty, with no error. [30s] works until a single scrape is missed. The rule is a range of at least four times the scrape interval: [1m] minimum with 15 second scrapes, [5m] as the usual choice. Dashboard tools can compute a safe value from the scrape interval for you.

Why must rate come before sum?

rate detects counter resets by seeing a value drop in a single series. If you sum counters across instances first, one instance restarting makes the total drop, which looks like a reset of the whole sum and gives a wrong rate. PromQL does not even allow rate(sum(x)[5m]) without a subquery. Always sum(rate(x[5m])): rate each counter, then aggregate the rates.

Why is increase() not a whole number for an integer counter?

Because it is an estimate. The first and last samples in the window are never exactly at its edges, so Prometheus extrapolates the observed increase to the full window length. The result, like 4871.297, is close to the true count but not exact. For small counts over short windows the error is relatively large; do not use increase for billing-grade exact numbers.

What is the difference between sum by and sum without?

sum by (uri) keeps only the listed labels in the result and sums everything else away, giving exactly one series per uri. sum without (instance) removes only the listed labels and keeps all others. without is more robust in recording rules and dashboards when new labels are added, since they are preserved rather than silently summed. by gives you a precise shape. The clause can go before or after the parentheses.

In an interview Mid

How would you calculate the request rate and error rate of a service in PromQL?

sum by (uri) (rate(http_server_requests_seconds_count{job="orders"}[5m]))

sum by (uri) (rate(http_server_requests_seconds_count{job="orders",status=~"5.."}[5m]))
  / sum by (uri) (rate(http_server_requests_seconds_count{job="orders"}[5m]))

Inside out: a selector (metric + label matchers), a range [5m], rate = per-second average increase, sum by (uri) = add up instances and statuses, one result per endpoint.

What makes it correct:

Also asked: Explain how rate() works and its edge cases. · Write the PromQL for CPU utilisation percentage per node and explain each part. · When would you use increase() instead of rate()?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.