How the numbers get into Prometheus
The orders app counts its requests in memory. Prometheus runs as a separate program. Before you can ask "how many 500s per second?", you need to know how the number travels from one to the other, where that is configured, and how to tell when it stopped travelling - because "the graph is flat" is often exactly that.
What you need to know already: time series, labels and samples (27.1); systemd units and systemctl reload (2.1, 2.21), SIGHUP (3.6), journalctl -u (2.30); ss -ltnp (9.5); HTTP GET and POST, status codes and curl (9.21); grep -E and jq including -r and @tsv (7.1, 7.11); YAML, the indentation-based config format of Docker Compose and Kubernetes manifests (11.31, 15.14).
The pull model
Prometheus pulls. Every scrape_interval (15 seconds on this box) it sends an HTTP GET to each target - a URL that serves metrics, usually ending in /metrics - and stores what comes back. One such request is a scrape. Applications do not push to it.
Why that matters:
- "Is it up?" is free. A failed scrape is itself data: Prometheus writes the series
up{job="...", instance="..."} 0(and1when the scrape worked). With push you cannot tell "no errors" from "the app died and stopped sending". - The app does not need to know where monitoring lives; Prometheus owns the list of targets.
- The cost is that Prometheus must be able to reach every target over the network, which is why in Kubernetes it runs inside the cluster. Short-lived batch jobs that end before a scrape can push to a helper called the Pushgateway, and that is the only place push belongs.
The exposition format
What a target returns is plain text, the exposition format. Look at this box's node exporter:
# the stack comes up with the next mission, where you run these yourself
curl -s localhost:9100/metrics | grep -E '^# (HELP|TYPE) node_cpu|^node_cpu_seconds_total\{cpu="0"'
# HELP node_cpu_seconds_total Seconds the CPUs spent in each mode.
# TYPE node_cpu_seconds_total counter
node_cpu_seconds_total{cpu="0",mode="idle"} 7475.689823714413
node_cpu_seconds_total{cpu="0",mode="iowait"} 32.939999999999735
node_cpu_seconds_total{cpu="0",mode="irq"} 0
node_cpu_seconds_total{cpu="0",mode="nice"} 0
node_cpu_seconds_total{cpu="0",mode="softirq"} 16.469999999999867
node_cpu_seconds_total{cpu="0",mode="steal"} 8.234999999999934
node_cpu_seconds_total{cpu="0",mode="system"} 205.875
node_cpu_seconds_total{cpu="0",mode="user"} 495.7901762855866
The command: curl -s fetches the URL silently (no progress bar); grep -E keeps only the HELP and TYPE comment lines of the CPU metric plus the series for CPU 0.
Reading the output: # HELP is the description, # TYPE the metric type (next section), then one line per series: name, labels in braces, value. These are the same modes top shows in its %Cpu(s) line (3.5): 7475 seconds idle, 495 in user code, 205 in the kernel. There is no timestamp: Prometheus stamps the sample with the scrape time. And there is no instance or job label: those are added by Prometheus from its own config, not by the target.
Large values print in exponent form (the Go library writes floats that way):
curl -s localhost:9100/metrics | grep ^node_memory_MemAvailable
node_memory_MemAvailable_bytes 3.72657440368519e+09
3.72657440368519e+09 is 3.73 x 10^9 bytes, about 3.7 GB: the same MemAvailable you read in /proc/meminfo (5.3). The Java client (Micrometer, in the orders app, 21.11) writes 84906.0 style numbers instead. Both are the same format to Prometheus.
The four metric types
Counter - only goes up, or goes back to 0 when the process restarts (a counter reset). Requests served, bytes sent, errors, CPU seconds. The name ends in _total (or _count, _sum, _bucket inside a histogram). You never graph a counter's raw value - "84906 requests since the JVM started" tells you nothing; you graph how fast it grows, its rate (27.8).
Gauge - a value that goes up and down. Memory in use, queue depth, temperature, connections in the pool, up. You graph it as it is.
Histogram - measurements (usually request durations) sorted into buckets by size, counted inside the application. A bucket is a counter of "how many requests took at most this long". One histogram produces three kinds of series:
curl -s localhost:8080/actuator/prometheus | grep 'uri="/api/checkout",le=' | head -10
http_server_requests_seconds_bucket{error="none",exception="none",method="POST",outcome="SUCCESS",status="200",uri="/api/checkout",le="0.025"} 1020.0
http_server_requests_seconds_bucket{...,le="0.05"} 2331.0
http_server_requests_seconds_bucket{...,le="0.1"} 4807.0
http_server_requests_seconds_bucket{...,le="0.25"} 8899.0
http_server_requests_seconds_bucket{...,le="0.5"} 10052.0
http_server_requests_seconds_bucket{...,le="1.0"} 10106.0
http_server_requests_seconds_bucket{...,le="2.5"} 10106.0
http_server_requests_seconds_bucket{...,le="5.0"} 10106.0
http_server_requests_seconds_bucket{...,le="10.0"} 10106.0
http_server_requests_seconds_bucket{...,le="+Inf"} 10106.0
http_server_requests_seconds_count{...} 10106.0
http_server_requests_seconds_sum{...} 1502.8
le means "less than or equal": le="0.25" counts requests that took at most 0.25 seconds. Buckets are cumulative: the 0.25 bucket (8899) includes everything in the 0.1 bucket (4807). So 4807 checkouts took under 100 ms, and 8899 - 4807 = 4092 took between 100 and 250 ms. +Inf (infinity) counts every request and equals _count. _sum is the total of all observed seconds, so _sum / _count is the average: 1502.8 / 10106 = 0.149 s. Each bucket is a counter, so like any counter you take its rate before doing anything else (27.15).
The buckets exist because the app was told which ones to keep; for Spring Boot that is management.metrics.distribution.slo.http.server.requests=25ms,50ms,... in /etc/orders/app.conf. Choose boundaries around your SLO threshold (0.8): if the SLO is "95% under 300 ms", a bucket at 0.3 makes the SLI exact.
Summary - percentiles computed inside the app, e.g. rpc_duration_seconds{quantile="0.99"} 0.21 ("the p99 is 210 ms"). Cheap to query, but you cannot combine them: there is no correct way to turn the p99 of pod A and the p99 of pod B into the p99 of the service (0.2 showed why percentiles do not average). Histograms can be combined - you add the buckets first. Prefer histograms for anything you run more than one copy of.
The scrape config on this box
/etc/prometheus/prometheus.yml is the Debian/Ubuntu default plus two jobs. A job is a named group of targets that do the same thing (both copies of orders are one job):
global:
scrape_interval: 15s
evaluation_interval: 15s
external_labels:
monitor: 'oncall-lab'
alerting:
alertmanagers:
- static_configs:
- targets: ['localhost:9093']
rule_files:
- "/etc/prometheus/rules/*.yml"
scrape_configs:
- job_name: 'prometheus'
scrape_interval: 5s
scrape_timeout: 5s
static_configs:
- targets: ['localhost:9090']
- job_name: node
static_configs:
- targets: ['localhost:9100']
- job_name: orders
metrics_path: /actuator/prometheus
static_configs:
- targets: ['localhost:8080', '10.0.3.21:8080']
labels:
team: orders
The top: global sets defaults for every job (evaluation_interval is how often rules run - 27.27). alerting says where to send alerts and rule_files where rule files live; both matter from 27.27 on.
scrape_configs is the list of jobs, and the lines worth knowing:
job_namebecomes thejoblabel on every series from that job.- each
targetsentry (host:port) becomes theinstancelabel. static_configsmeans "this exact, fixed list of targets" (27.4 shows lists that change by themselves).labelsunder a static config are added to every series of those targets.metrics_pathdefaults to/metrics. Spring Boot serves/actuator/prometheus, so forgetting this line givesup == 0withserver returned HTTP status 404 Not Found.scrape_intervalper job overrides the global one;scrape_timeoutis how long one scrape may take before it counts as failed. Keep the interval the same across jobs unless you have a reason: mixed intervals make rate windows confusing (27.8).
Targets and health
Prometheus has an HTTP API on port 9090. /api/v1/targets returns every target and whether its last scrape worked, as JSON:
curl -s localhost:9090/api/v1/targets | jq -r '.data.activeTargets[] | [.labels.job, .labels.instance, .health, .lastError] | @tsv'
prometheus localhost:9090 up
node localhost:9100 up
orders localhost:8080 up
orders 10.0.3.21:8080 up
lab localhost:9199 up
The jq program walks the activeTargets array and prints four fields per target as tab-separated text (-r for raw strings, @tsv for tabs, 7.11). All five are up, so lastError is empty. When a target is down, lastError says why, in the words of Go (the language Prometheus is written in):
Get "http://localhost:9100/metrics": dial tcp 127.0.0.1:9100: connect: connection refused
server returned HTTP status 404 Not Found
Get "http://10.0.3.99:9100/metrics": context deadline exceeded
Get "http://nodes.lab:9100/metrics": dial tcp: lookup nodes.lab on 127.0.0.53:53: no such host
Those map one to one onto the networking chapters: refused (nothing listening, 9.1), wrong path (404, 9.21), timeout - "context deadline exceeded" means the scrape took longer than scrape_timeout, usually dropped packets (9.1) - and a DNS name that does not resolve (8.16).
The TSDB in one paragraph
New samples land in the head block: the most recent two hours, kept in memory (plus a write-ahead log on disk in /var/lib/prometheus/metrics2/wal, so a crash loses nothing). Every two hours the head is written out as an unchangeable block directory on disk. Retention - how long data is kept - is 15 days by default. Memory use is driven by the number of active series in the head, which Prometheus reports about itself as prometheus_tsdb_head_series: that is the number to look at when Prometheus starts eating RAM.
Lookback and staleness
When you ask for a series' value "now", Prometheus returns its most recent sample - but only if that sample is not older than the lookback delta (5 minutes by default). When a scrape fails or a target disappears, Prometheus writes a staleness marker (a special "this series ended" sample), and the series vanishes from queries immediately instead of lingering for five minutes. Two consequences you will meet:
- after you stop an exporter, its series (
node_load1...) disappear right away, butupfor that target stays and becomes0(Prometheus writesupitself). - after you remove a target from the config, even
updisappears. A check forup == 0then finds nothing at all - no series, no zero. 27.15 shows the function built for "this should exist and does not".
Reloading, and the one gotcha of the Ubuntu package
After editing prometheus.yml you must tell Prometheus to re-read it. Prometheus has an HTTP endpoint for that, /-/reload, part of what it calls the lifecycle API. The package starts Prometheus with ARGS="" from /etc/default/prometheus, so the flag that enables it, --web.enable-lifecycle, is off:
curl -s localhost:9090/-/reload -X POST
Lifecycle API is not enabled.
(-X POST makes curl send a POST instead of a GET; the endpoint only accepts POST.) Reload with sudo systemctl reload prometheus instead: the unit's ExecReload=/bin/kill -HUP $MAINPID sends SIGHUP (3.6), and Prometheus reads its config again on SIGHUP.
Always run promtool check config /etc/prometheus/prometheus.yml first. promtool is the checking and querying tool that ships with Prometheus; check config parses the file and every rule file it points to, and prints SUCCESS or the error. A reload with a broken file is rejected and the old configuration keeps running - which is safe, and also the reason people think their change is live when it is not. The journal says so:
$ journalctl -u prometheus -n 2 --no-pager
... level=error msg="Error reloading config" err="couldn't load configuration (--config.file=\"/etc/prometheus/prometheus.yml\"): parsing YAML file /etc/prometheus/prometheus.yml: yaml: unmarshal errors: line 57: field scrape_intervall not found in type config.plain"
... systemd[1]: Reloaded prometheus.service - Monitoring system and time series database.
The two lines: Prometheus says it could not load the file (line 57 has a field scrape_intervall, with a typo, that it does not know), and then systemd happily reports "Reloaded". systemd only sent the signal; it has no idea whether Prometheus accepted the file. The metric prometheus_config_last_reload_successful does (1 = yes, 0 = no), and is worth an alert of its own.
Querying from a terminal
Prometheus's query language is PromQL (27.8 teaches it). The simplest query is a metric name, and there are three ways to send one from the shell, all hitting the same HTTP API:
promtool query instant http://localhost:9090 'up'
curl -s localhost:9090/api/v1/query --data-urlencode 'query=up' | jq
curl -sG localhost:9090/api/v1/query --data-urlencode 'query=up' | jq
promtool query instant <server> '<query>' asks for the value now and prints one line per series - the easiest to read. The curl versions return JSON: --data-urlencode 'query=up' sends the query as a form field, safely encoded; alone it makes a POST, and -G turns it into a GET with the query in the URL. Both work.
Do not paste PromQL raw into a URL. curl treats [ and { as its own pattern syntax:
$ curl 'localhost:9090/api/v1/query?query=rate(up[5m])'
curl: (3) bad range in URL position 42:
localhost:9090/api/v1/query?query=rate(up[5m])
^
and {job="node"} silently loses its braces, so Prometheus receives upjob="node" and answers with a parse error. --data-urlencode fixes both.
What you can now do
- Explain the pull model and why
upis free, and read exposition text: HELP, TYPE, labels, value. - Tell counter, gauge, histogram and summary apart, and read cumulative buckets.
- Read prometheus.yml, check a target's health and lastError, and reload safely:
promtool check config,systemctl reload, then the journal.