OnCallReady

Lesson 27.2 · Observability I: Prometheus & PromQL · 24 min read

Prometheus: the pull model, exposition, the four metric types

In plain words

Imagine a school nurse who walks around every 15 minutes with a clipboard, visits each classroom and writes down what the notice on the door says: "kids present: 24, kids who sneezed today: 7". If a classroom door is locked, she notes "no answer", which itself tells her something is wrong. The classrooms never have to run to her.

Prometheus is that nurse. Every scrape_interval it does an HTTP GET on each target's /metrics (for orders, /actuator/prometheus) and stores what comes back, adding job and instance labels. A failed visit becomes up 0. The notices hold four kinds of numbers: counters that only go up, gauges that go up and down, histograms of durations in buckets, and summaries.

How the numbers get into Prometheus

The orders app counts its requests in memory. Prometheus runs as a separate program. Before you can ask "how many 500s per second?", you need to know how the number travels from one to the other, where that is configured, and how to tell when it stopped travelling - because "the graph is flat" is often exactly that.

What you need to know already: time series, labels and samples (27.1); systemd units and systemctl reload (2.1, 2.21), SIGHUP (3.6), journalctl -u (2.30); ss -ltnp (9.5); HTTP GET and POST, status codes and curl (9.21); grep -E and jq including -r and @tsv (7.1, 7.11); YAML, the indentation-based config format of Docker Compose and Kubernetes manifests (11.31, 15.14).

The pull model

Prometheus pulls. Every scrape_interval (15 seconds on this box) it sends an HTTP GET to each target - a URL that serves metrics, usually ending in /metrics - and stores what comes back. One such request is a scrape. Applications do not push to it.

Why that matters:

The exposition format

What a target returns is plain text, the exposition format. Look at this box's node exporter:

# the stack comes up with the next mission, where you run these yourself
curl -s localhost:9100/metrics | grep -E '^# (HELP|TYPE) node_cpu|^node_cpu_seconds_total\{cpu="0"'
# HELP node_cpu_seconds_total Seconds the CPUs spent in each mode.
# TYPE node_cpu_seconds_total counter
node_cpu_seconds_total{cpu="0",mode="idle"} 7475.689823714413
node_cpu_seconds_total{cpu="0",mode="iowait"} 32.939999999999735
node_cpu_seconds_total{cpu="0",mode="irq"} 0
node_cpu_seconds_total{cpu="0",mode="nice"} 0
node_cpu_seconds_total{cpu="0",mode="softirq"} 16.469999999999867
node_cpu_seconds_total{cpu="0",mode="steal"} 8.234999999999934
node_cpu_seconds_total{cpu="0",mode="system"} 205.875
node_cpu_seconds_total{cpu="0",mode="user"} 495.7901762855866

The command: curl -s fetches the URL silently (no progress bar); grep -E keeps only the HELP and TYPE comment lines of the CPU metric plus the series for CPU 0.

Reading the output: # HELP is the description, # TYPE the metric type (next section), then one line per series: name, labels in braces, value. These are the same modes top shows in its %Cpu(s) line (3.5): 7475 seconds idle, 495 in user code, 205 in the kernel. There is no timestamp: Prometheus stamps the sample with the scrape time. And there is no instance or job label: those are added by Prometheus from its own config, not by the target.

Large values print in exponent form (the Go library writes floats that way):

curl -s localhost:9100/metrics | grep ^node_memory_MemAvailable
node_memory_MemAvailable_bytes 3.72657440368519e+09

3.72657440368519e+09 is 3.73 x 10^9 bytes, about 3.7 GB: the same MemAvailable you read in /proc/meminfo (5.3). The Java client (Micrometer, in the orders app, 21.11) writes 84906.0 style numbers instead. Both are the same format to Prometheus.

The four metric types

Counter - only goes up, or goes back to 0 when the process restarts (a counter reset). Requests served, bytes sent, errors, CPU seconds. The name ends in _total (or _count, _sum, _bucket inside a histogram). You never graph a counter's raw value - "84906 requests since the JVM started" tells you nothing; you graph how fast it grows, its rate (27.8).

Gauge - a value that goes up and down. Memory in use, queue depth, temperature, connections in the pool, up. You graph it as it is.

Histogram - measurements (usually request durations) sorted into buckets by size, counted inside the application. A bucket is a counter of "how many requests took at most this long". One histogram produces three kinds of series:

curl -s localhost:8080/actuator/prometheus | grep 'uri="/api/checkout",le=' | head -10
http_server_requests_seconds_bucket{error="none",exception="none",method="POST",outcome="SUCCESS",status="200",uri="/api/checkout",le="0.025"} 1020.0
http_server_requests_seconds_bucket{...,le="0.05"} 2331.0
http_server_requests_seconds_bucket{...,le="0.1"} 4807.0
http_server_requests_seconds_bucket{...,le="0.25"} 8899.0
http_server_requests_seconds_bucket{...,le="0.5"} 10052.0
http_server_requests_seconds_bucket{...,le="1.0"} 10106.0
http_server_requests_seconds_bucket{...,le="2.5"} 10106.0
http_server_requests_seconds_bucket{...,le="5.0"} 10106.0
http_server_requests_seconds_bucket{...,le="10.0"} 10106.0
http_server_requests_seconds_bucket{...,le="+Inf"} 10106.0
http_server_requests_seconds_count{...} 10106.0
http_server_requests_seconds_sum{...} 1502.8

le means "less than or equal": le="0.25" counts requests that took at most 0.25 seconds. Buckets are cumulative: the 0.25 bucket (8899) includes everything in the 0.1 bucket (4807). So 4807 checkouts took under 100 ms, and 8899 - 4807 = 4092 took between 100 and 250 ms. +Inf (infinity) counts every request and equals _count. _sum is the total of all observed seconds, so _sum / _count is the average: 1502.8 / 10106 = 0.149 s. Each bucket is a counter, so like any counter you take its rate before doing anything else (27.15).

The buckets exist because the app was told which ones to keep; for Spring Boot that is management.metrics.distribution.slo.http.server.requests=25ms,50ms,... in /etc/orders/app.conf. Choose boundaries around your SLO threshold (0.8): if the SLO is "95% under 300 ms", a bucket at 0.3 makes the SLI exact.

Summary - percentiles computed inside the app, e.g. rpc_duration_seconds{quantile="0.99"} 0.21 ("the p99 is 210 ms"). Cheap to query, but you cannot combine them: there is no correct way to turn the p99 of pod A and the p99 of pod B into the p99 of the service (0.2 showed why percentiles do not average). Histograms can be combined - you add the buckets first. Prefer histograms for anything you run more than one copy of.

The scrape config on this box

/etc/prometheus/prometheus.yml is the Debian/Ubuntu default plus two jobs. A job is a named group of targets that do the same thing (both copies of orders are one job):

global:
  scrape_interval:     15s
  evaluation_interval: 15s
  external_labels:
      monitor: 'oncall-lab'

alerting:
  alertmanagers:
  - static_configs:
    - targets: ['localhost:9093']

rule_files:
  - "/etc/prometheus/rules/*.yml"

scrape_configs:
  - job_name: 'prometheus'
    scrape_interval: 5s
    scrape_timeout: 5s
    static_configs:
      - targets: ['localhost:9090']

  - job_name: node
    static_configs:
      - targets: ['localhost:9100']

  - job_name: orders
    metrics_path: /actuator/prometheus
    static_configs:
      - targets: ['localhost:8080', '10.0.3.21:8080']
        labels:
          team: orders

The top: global sets defaults for every job (evaluation_interval is how often rules run - 27.27). alerting says where to send alerts and rule_files where rule files live; both matter from 27.27 on.

scrape_configs is the list of jobs, and the lines worth knowing:

Targets and health

Prometheus has an HTTP API on port 9090. /api/v1/targets returns every target and whether its last scrape worked, as JSON:

curl -s localhost:9090/api/v1/targets | jq -r '.data.activeTargets[] | [.labels.job, .labels.instance, .health, .lastError] | @tsv'
prometheus  localhost:9090   up
node        localhost:9100   up
orders      localhost:8080   up
orders      10.0.3.21:8080   up
lab         localhost:9199   up

The jq program walks the activeTargets array and prints four fields per target as tab-separated text (-r for raw strings, @tsv for tabs, 7.11). All five are up, so lastError is empty. When a target is down, lastError says why, in the words of Go (the language Prometheus is written in):

Get "http://localhost:9100/metrics": dial tcp 127.0.0.1:9100: connect: connection refused
server returned HTTP status 404 Not Found
Get "http://10.0.3.99:9100/metrics": context deadline exceeded
Get "http://nodes.lab:9100/metrics": dial tcp: lookup nodes.lab on 127.0.0.53:53: no such host

Those map one to one onto the networking chapters: refused (nothing listening, 9.1), wrong path (404, 9.21), timeout - "context deadline exceeded" means the scrape took longer than scrape_timeout, usually dropped packets (9.1) - and a DNS name that does not resolve (8.16).

The TSDB in one paragraph

New samples land in the head block: the most recent two hours, kept in memory (plus a write-ahead log on disk in /var/lib/prometheus/metrics2/wal, so a crash loses nothing). Every two hours the head is written out as an unchangeable block directory on disk. Retention - how long data is kept - is 15 days by default. Memory use is driven by the number of active series in the head, which Prometheus reports about itself as prometheus_tsdb_head_series: that is the number to look at when Prometheus starts eating RAM.

Lookback and staleness

When you ask for a series' value "now", Prometheus returns its most recent sample - but only if that sample is not older than the lookback delta (5 minutes by default). When a scrape fails or a target disappears, Prometheus writes a staleness marker (a special "this series ended" sample), and the series vanishes from queries immediately instead of lingering for five minutes. Two consequences you will meet:

Reloading, and the one gotcha of the Ubuntu package

After editing prometheus.yml you must tell Prometheus to re-read it. Prometheus has an HTTP endpoint for that, /-/reload, part of what it calls the lifecycle API. The package starts Prometheus with ARGS="" from /etc/default/prometheus, so the flag that enables it, --web.enable-lifecycle, is off:

curl -s localhost:9090/-/reload -X POST
Lifecycle API is not enabled.

(-X POST makes curl send a POST instead of a GET; the endpoint only accepts POST.) Reload with sudo systemctl reload prometheus instead: the unit's ExecReload=/bin/kill -HUP $MAINPID sends SIGHUP (3.6), and Prometheus reads its config again on SIGHUP.

Always run promtool check config /etc/prometheus/prometheus.yml first. promtool is the checking and querying tool that ships with Prometheus; check config parses the file and every rule file it points to, and prints SUCCESS or the error. A reload with a broken file is rejected and the old configuration keeps running - which is safe, and also the reason people think their change is live when it is not. The journal says so:

$ journalctl -u prometheus -n 2 --no-pager
... level=error msg="Error reloading config" err="couldn't load configuration (--config.file=\"/etc/prometheus/prometheus.yml\"): parsing YAML file /etc/prometheus/prometheus.yml: yaml: unmarshal errors:   line 57: field scrape_intervall not found in type config.plain"
... systemd[1]: Reloaded prometheus.service - Monitoring system and time series database.

The two lines: Prometheus says it could not load the file (line 57 has a field scrape_intervall, with a typo, that it does not know), and then systemd happily reports "Reloaded". systemd only sent the signal; it has no idea whether Prometheus accepted the file. The metric prometheus_config_last_reload_successful does (1 = yes, 0 = no), and is worth an alert of its own.

Querying from a terminal

Prometheus's query language is PromQL (27.8 teaches it). The simplest query is a metric name, and there are three ways to send one from the shell, all hitting the same HTTP API:

promtool query instant http://localhost:9090 'up'
curl -s localhost:9090/api/v1/query --data-urlencode 'query=up' | jq
curl -sG localhost:9090/api/v1/query --data-urlencode 'query=up' | jq

promtool query instant <server> '<query>' asks for the value now and prints one line per series - the easiest to read. The curl versions return JSON: --data-urlencode 'query=up' sends the query as a form field, safely encoded; alone it makes a POST, and -G turns it into a GET with the query in the URL. Both work.

Do not paste PromQL raw into a URL. curl treats [ and { as its own pattern syntax:

$ curl 'localhost:9090/api/v1/query?query=rate(up[5m])'
curl: (3) bad range in URL position 42:
localhost:9090/api/v1/query?query=rate(up[5m])
                                         ^

and {job="node"} silently loses its braces, so Prometheus receives upjob="node" and answers with a parse error. --data-urlencode fixes both.

What you can now do

Why it helps

When a target shows up == 0, the lastError in /api/v1/targets tells you which of four problems it is: connection refused, 404 Not Found (the classic missing metrics_path: /actuator/prometheus), context deadline exceeded, or DNS no such host. That is minutes instead of an hour.

Knowing the metric types stops you making the most common dashboard mistakes: graphing a counter's raw value, trying to aggregate summaries across pods, choosing histogram buckets that make your SLO threshold unmeasurable. And the Ubuntu gotchas are real on your VM: the lifecycle API is off, systemctl reload says "Reloaded" even when Prometheus rejected the file, and only prometheus_config_last_reload_successful or the journal tells the truth. "Explain the four metric types" is a classic interview question.

Commands in this lesson

journalctl curl

FAQ

What is the difference between a histogram and a summary?

Both measure distributions, usually durations. A histogram counts observations into fixed buckets in the app and exposes them as counters; Prometheus computes quantiles at query time with histogram_quantile, and buckets from many instances can be summed first, so it aggregates correctly. A summary computes quantiles inside the app over a sliding window; they are cheap to query but cannot be combined across instances. Prefer histograms for anything with more than one replica.

Why is there no job or instance label in the /metrics output?

Because Prometheus adds them itself, from the scrape configuration: job from job_name, instance from the target address, plus any labels: of the static config. The target does not know how it is being scraped. If a target does expose its own job or instance label, Prometheus renames it to exported_job or exported_instance unless honor_labels: true is set.

Why does my Prometheus reload seem to do nothing?

If the new file has an error, Prometheus rejects the reload and keeps the old configuration running, while systemctl reload still prints "Reloaded" because systemd only sent SIGHUP. Check journalctl -u prometheus for Error reloading config, and the metric prometheus_config_last_reload_successful. Always run promtool check config before reloading. On Ubuntu, /-/reload over HTTP is disabled because --web.enable-lifecycle is off.

Why do values like 3.72657440368519e+09 appear?

The Go client libraries format floats with the shortest representation, which uses exponent notation for large numbers, so 3,726,574,403 bytes prints as 3.72657440368519e+09. The Java client writes 84906.0-style doubles. It is all the same text exposition format to Prometheus: every sample value is a 64-bit float. Tools like jq display them normally.

How do I query Prometheus from the terminal without breaking the query?

Use promtool query instant http://localhost:9090 'expr', or curl with --data-urlencode 'query=expr' against /api/v1/query. Do not paste PromQL raw into a URL: curl treats [] and {} as its own globbing syntax, so rate(up[5m]) fails with bad range in URL, and braces in selectors silently disappear, giving a parse error from Prometheus.

In an interview Mid

Explain the four Prometheus metric types.

How the numbers arrive: Prometheus pulls - it scrapes each target's /metrics every scrape_interval and stores the samples, adding job and instance. A failed scrape writes up 0, so "is it up?" comes for free; the target's lastError says why.

Also asked: A Prometheus target is down. How do you troubleshoot it? · Why does Prometheus pull metrics instead of having applications push them? · How do you choose histogram buckets?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.