Why Prometheus has to call you
In Ch 27 you learned to ask Prometheus questions: "how many requests failed in the last five minutes?". But nobody sits in front of a query at 3 a.m. If the node exporter dies at night, the graph shows a gap and nobody sees it until morning. You need Prometheus to ask the question for you, every few seconds, and shout when the answer is bad. That is an alerting rule.
What you need to know already: what up is and how a scrape fails (27.2); selectors, rate() and sum by (27.8); comparison operators as filters, bool, and absent() (27.15); recording rules, rule files, rule_files and promtool check rules (27.27); page vs ticket, symptom vs cause, alert fatigue and runbooks (0.21); curl -s (9.21); jq -c and jq -r (7.11); writing a root-owned file with sudo tee and a quoted heredoc (6.18); systemctl reload (2.21).
The words you need first
- Alerting rule - a PromQL expression plus a name, stored in a rule file. Prometheus runs it again and again and turns its result into alerts.
- Evaluation - one run of the rule. It happens every
evaluation_interval(a setting in prometheus.yml; 15 s on this box, 27.2). - Alert instance - one series in the result of an alerting rule. The rule is the recipe; each instance is one "thing that is wrong right now", told apart by its labels.
- Annotation - free text attached to an alert (a summary, a link). Unlike a label, it does not identify the alert; it is the message a human reads.
An alert is a query that returns something
This is the whole model, and everything in this chapter builds on it: every series in the result is an alert. An empty result means everything is fine. That is why the comparison operators from 27.15 are the heart of alert expressions: without bool, up == 0 is a filter that keeps only the series that are 0, so it returns exactly the dead targets and nothing else.
A rule file with one alerting rule:
groups:
- name: targets
rules:
- alert: TargetDown
expr: up == 0
for: 2m
labels:
severity: warning
annotations:
summary: "{{ $labels.job }} target {{ $labels.instance }} is down"
description: "Prometheus cannot scrape {{ $labels.instance }} ({{ $labels.job }}) for 2 minutes."
runbook_url: "https://wiki.lab/runbooks/target-down"
groups and rules are the same structure you used for recording rules in 27.27; a group can hold both kinds. Field by field:
alert- the alert's name. Prometheus adds it to every instance as the labelalertname, soalertname="TargetDown"is how you select it later.expr- the query. Each series it returns is one alert instance.up == 0with two dead targets is two alerts, one perinstancelabel.for- how long the condition must be continuously true before the alert counts. Until then the alert is pending (waiting, not yet sent anywhere).labels- extra labels added to every instance.severityis the convention for "how urgent": this lab usespage(wake someone),warningandticket(someone looks during work hours), the same split as 0.21. Keep label values static: a label that contains the current value ({{ $value }}) changes at every evaluation, so Prometheus sees a brand-new alert each time andfornever completes.annotations- the human text. They are templates - text with placeholders filled in at evaluation time, the same Go template syntax as Helm (25.24).{{ $labels.job }}becomes the alert'sjoblabel,{{ $value }}its value. Helper functions format numbers:humanize(1234567 -> 1.235M),humanizePercentage(0.0642 -> 6.42%),humanizeDuration(seconds -> "2m 3s"),printf "%.2f"(two decimals).
runbook_url is just a conventional annotation name: the link to the page that tells the person paged what to do (0.21).
inactive, pending, firing
Every alert instance is in one of three states:
- inactive - the expression does not return this series. Nothing is wrong.
- pending - the expression returns it, but not yet for the whole
for. - firing - it has been returned continuously for at least
for. Only now is it sent on.
You can watch the states through Prometheus's HTTP API, the same web API promtool query talks to. /api/v1/alerts lists every pending or firing alert as JSON:
# with the node exporter stopped (the next mission stops it)
$ curl -s localhost:9090/api/v1/alerts | jq -c '.data.alerts[] | [.labels.alertname, .state, .labels.instance]'
["TargetDown","pending","localhost:9100"]
... one minute later ...
["TargetDown","firing","localhost:9100"]
curl -s fetches the URL without a progress bar (9.21). The jq program walks into .data.alerts, and for each alert ([]) builds a small array of its name, state and instance; -c prints each array on one compact line. The output reads: the localhost:9100 target is pending, and a minute later it is firing.
What Prometheus does at each evaluation, per alert instance:
condition true, was inactive -> pending (activeAt = now) [firing at once if for: 0]
condition true, was pending -> firing once now - activeAt >= for
condition false -> resolved, gone
activeAt is the moment the condition first became true: the start of the for clock.
The alert only leaves Prometheus when it fires. Pending alerts exist only inside Prometheus. The program that receives firing alerts and turns them into chat messages and pages is Alertmanager; this chapter's third lesson (28.7) is about it. For now: Prometheus fires, Alertmanager notifies.
Prometheus also writes every pending and firing alert into a real time series called ALERTS, which you can query like any metric. That is how you find out, afterwards, whether an alert did or did not fire:
$ promtool query instant http://localhost:9090 'ALERTS'
ALERTS{alertname="TargetDown", alertstate="firing", instance="localhost:9100", job="node", severity="warning"} => 1
The series carries the alert's labels plus alertstate (pending or firing), with the value 1 while it is active. max_over_time(ALERTS{alertname="X"}[3h]) - the highest value each matching series had in the last three hours (27.22) - tells you whether X was pending or firing at any point in that time. You will use exactly that in this chapter's first incident (28.13).
Choosing for:
for trades detection speed for noise. With for: 5m a blip of 30 seconds never pages; a real outage pages five minutes late. Guidelines:
- Never longer than the outages you care about.
for: 1hon an error-rate alert means a 45-minute outage never pages. It sounds absurd written down; it is in production configs everywhere. forresets on one false evaluation. A flapping condition (true, false, true... around the threshold, 0.21) withfor: 10mcan stay pending forever. Smooth the expression instead: a longer rate window, oravg_over_time(the average of a series over a window, likemax_over_timebut the mean), or usekeep_firing_for.- For SLO burn-rate alerts the windows do the smoothing and
foris short or absent (the SLO lesson later in this chapter, 28.14). keep_firing_for: 10m(Prometheus 2.42 and later) keeps a firing alert firing for that long after the condition clears. It stops "resolved / firing again" storms from a noisy condition.
A Prometheus restart does not reset a pending for clock: Prometheus stores the clocks in a series called ALERTS_FOR_STATE and restores them on startup, as long as it was down for less than --rules.alert.for-outage-tolerance (a command-line flag, 1h by default).
Symptoms, not causes
Chapter 0 made the argument (0.21); this is what it looks like in a rule file.
# cause-based: fires often, means little on its own
- alert: NodeHighCPU
expr: 100 * (1 - avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m]))) > 80
for: 5m
labels: { severity: warning }
# symptom-based: fires when users are affected, whatever the cause
- alert: OrdersErrorBudgetBurn
expr: job:slo_errors_per_request:ratio_rate1h > (14.4 * 0.001) and job:slo_errors_per_request:ratio_rate5m > (14.4 * 0.001)
labels: { severity: page }
The first rule is the CPU query from 27.8 (100 minus the idle percentage) with > 80 on the end: it returns every instance above 80% busy. The second reads two recorded error ratios; the numbers 14.4 and 0.001 are burn-rate arithmetic from 0.17, built step by step in 28.14. For now read it as "orders is failing fast enough to spend its error budget, both over the last hour and right now".
High CPU during a batch job hurts nobody; a database failover that errors 8% of checkouts may not move CPU at all. Page on symptoms (errors, latency, availability as users see them). Keep cause alerts, but as tickets or dashboard annotations at low severity - they help diagnosis, they must not wake anyone.
Alert fatigue (0.21) comes from: pages nobody acts on, duplicate alerts for one problem (fixed with grouping and inhibition in Alertmanager, 28.7), flapping alerts, thresholds copied from a blog post, and alerts without a runbook. Every page should be urgent, actionable, and about users. If an alert fires and the right response is "ignore it", delete the alert.
Dead targets: up == 0 and absent()
Two different failures, two different expressions:
- alert: TargetDown # the target exists but cannot be scraped
expr: up == 0
for: 2m
- alert: NodeExporterAbsent # the target is gone from service discovery
expr: absent(up{job="node"})
for: 5m
up == 0 needs an up series to exist. When a target disappears from the config entirely (a relabel rule that drops it, a ServiceMonitor with the wrong selector - both 27.4 - or a job someone renamed), there is no up series at all, the expression returns nothing, and the alert is silent exactly when monitoring is broken. absent() (27.15) turns "nothing" into one series with value 1, so it fires on exactly that. Write one absent(up{job="..."}) per job you cannot live without, and the same for key metrics: absent(http_server_requests_seconds_count{job="orders"}) catches the day an upgrade renames the metric.
There is also absent_over_time(x[10m]): it returns 1 if x had no samples at all in the last 10 minutes - handy for series that only appear now and then.
Disk will fill, and other predictions
- alert: DiskWillFillIn4Hours
expr: |
predict_linear(node_filesystem_avail_bytes{fstype!~"tmpfs|overlay"}[6h], 4 * 3600) < 0
and
node_filesystem_avail_bytes{fstype!~"tmpfs|overlay"} / node_filesystem_size_bytes < 0.2
for: 30m
labels:
severity: ticket
annotations:
summary: "{{ $labels.mountpoint }} on {{ $labels.instance }} will be full within 4 hours"
Part by part: predict_linear(x[6h], 4 * 3600) (27.15) draws a straight line through the last 6 hours of free space and says where it will be in 4 hours (4 x 3600 seconds); < 0 keeps the disks predicted to run out. The fstype!~"tmpfs|overlay" matcher leaves out memory-backed and container filesystems. expr: | is YAML for "the next indented lines are one string", so a long expression can span lines.
The and clause keeps a left-hand series only if a series with the same labels also exists on the right. Here the right side is "less than 20% free", so the alert stays quiet on a disk that is 90% empty but briefly filling fast (a big download). and, or and unless match on labels like the arithmetic operators (27.15); both sides here have the same labels, so it just works.
Annotations that help at 3 a.m.
annotations:
summary: "orders 5xx ratio is {{ $value | humanizePercentage }}"
description: "{{ $labels.job }} has served more than 1.44% errors for 1 hour. Dashboard: https://dashboards.lab/d/orders"
runbook_url: "https://wiki.lab/runbooks/orders-errors"
{{ $value }} is the value of the result series; for a comparison filter like x > 0.05 that is the left-hand side, x. The | passes it through humanizePercentage, which turns 0.0642 into 6.42%. A link to the dashboard and the runbook turns a page into a starting point.
Loading and checking
Rule files are listed under rule_files in prometheus.yml; this box uses a glob (a * pattern), so any /etc/prometheus/rules/*.yml is picked up on reload:
$ promtool check rules /etc/prometheus/rules/targets.yml
Checking /etc/prometheus/rules/targets.yml
SUCCESS: 2 rules found
$ sudo systemctl reload prometheus
$ curl -s localhost:9090/api/v1/rules | jq -c '.data.groups[].rules[] | [.name, .state, .health]'
["NodeHighCPU","inactive","ok"]
["OrdersHighErrorRate","inactive","ok"]
["TargetDown","inactive","ok"]
["NodeExporterAbsent","inactive","ok"]
promtool check rules parses the file and every expression in it without running anything: SUCCESS: 2 rules found means the YAML and the PromQL are valid. The reload makes Prometheus read the files again. /api/v1/rules lists every loaded rule; the jq program prints, per rule, its name, its state (inactive, pending or firing - of the rule as a whole) and its health (ok, or err if the last evaluation failed).
"inactive, ok" is the healthy state of an alert that is not firing. It is also the state of an alert whose expression can never return anything - a typo in a label value, a regex that matches no series. Prometheus cannot tell a healthy system from a broken rule. Only a test can (next lesson, 28.4), and the second incident of this chapter (28.23) is exactly that.
What you can now do
- Write an alerting rule with
expr,for, static labels and templated annotations, check it with promtool and load it. - Follow an alert through inactive, pending and firing in
/api/v1/alerts, and find its history inALERTS. - Choose between
up == 0andabsent(), and between a symptom page and a cause ticket.