OnCallReady

Lesson 28.1 · Observability II: Alerting, Alertmanager & SLO alerts · 24 min read

Alerting rules: for, labels, symptoms, dead targets

In plain words

Imagine a baby monitor with a rule: "if the baby cries for more than two minutes, wake the parents". A single hiccup does not wake anyone; the light turns yellow, waiting. If the crying continues for two full minutes, the light turns red and the parents' phone rings. If the baby stops for a second, the timer starts over.

A Prometheus alerting rule is that monitor. The expr is the question, like up == 0; every series in the result is one alert. for: 2m is the waiting time: the alert is pending until the condition has been true continuously for two minutes, then firing. labels like severity decide who gets called, annotations are the message. And absent() notices when the baby monitor itself was unplugged.

Why Prometheus has to call you

In Ch 27 you learned to ask Prometheus questions: "how many requests failed in the last five minutes?". But nobody sits in front of a query at 3 a.m. If the node exporter dies at night, the graph shows a gap and nobody sees it until morning. You need Prometheus to ask the question for you, every few seconds, and shout when the answer is bad. That is an alerting rule.

What you need to know already: what up is and how a scrape fails (27.2); selectors, rate() and sum by (27.8); comparison operators as filters, bool, and absent() (27.15); recording rules, rule files, rule_files and promtool check rules (27.27); page vs ticket, symptom vs cause, alert fatigue and runbooks (0.21); curl -s (9.21); jq -c and jq -r (7.11); writing a root-owned file with sudo tee and a quoted heredoc (6.18); systemctl reload (2.21).

The words you need first

An alert is a query that returns something

This is the whole model, and everything in this chapter builds on it: every series in the result is an alert. An empty result means everything is fine. That is why the comparison operators from 27.15 are the heart of alert expressions: without bool, up == 0 is a filter that keeps only the series that are 0, so it returns exactly the dead targets and nothing else.

A rule file with one alerting rule:

groups:
  - name: targets
    rules:
      - alert: TargetDown
        expr: up == 0
        for: 2m
        labels:
          severity: warning
        annotations:
          summary: "{{ $labels.job }} target {{ $labels.instance }} is down"
          description: "Prometheus cannot scrape {{ $labels.instance }} ({{ $labels.job }}) for 2 minutes."
          runbook_url: "https://wiki.lab/runbooks/target-down"

groups and rules are the same structure you used for recording rules in 27.27; a group can hold both kinds. Field by field:

runbook_url is just a conventional annotation name: the link to the page that tells the person paged what to do (0.21).

inactive, pending, firing

Every alert instance is in one of three states:

You can watch the states through Prometheus's HTTP API, the same web API promtool query talks to. /api/v1/alerts lists every pending or firing alert as JSON:

# with the node exporter stopped (the next mission stops it)
$ curl -s localhost:9090/api/v1/alerts | jq -c '.data.alerts[] | [.labels.alertname, .state, .labels.instance]'
["TargetDown","pending","localhost:9100"]
   ... one minute later ...
["TargetDown","firing","localhost:9100"]

curl -s fetches the URL without a progress bar (9.21). The jq program walks into .data.alerts, and for each alert ([]) builds a small array of its name, state and instance; -c prints each array on one compact line. The output reads: the localhost:9100 target is pending, and a minute later it is firing.

What Prometheus does at each evaluation, per alert instance:

condition true, was inactive    -> pending (activeAt = now)   [firing at once if for: 0]
condition true, was pending     -> firing once now - activeAt >= for
condition false                 -> resolved, gone

activeAt is the moment the condition first became true: the start of the for clock.

The alert only leaves Prometheus when it fires. Pending alerts exist only inside Prometheus. The program that receives firing alerts and turns them into chat messages and pages is Alertmanager; this chapter's third lesson (28.7) is about it. For now: Prometheus fires, Alertmanager notifies.

Prometheus also writes every pending and firing alert into a real time series called ALERTS, which you can query like any metric. That is how you find out, afterwards, whether an alert did or did not fire:

$ promtool query instant http://localhost:9090 'ALERTS'
ALERTS{alertname="TargetDown", alertstate="firing", instance="localhost:9100", job="node", severity="warning"} => 1

The series carries the alert's labels plus alertstate (pending or firing), with the value 1 while it is active. max_over_time(ALERTS{alertname="X"}[3h]) - the highest value each matching series had in the last three hours (27.22) - tells you whether X was pending or firing at any point in that time. You will use exactly that in this chapter's first incident (28.13).

Choosing for:

for trades detection speed for noise. With for: 5m a blip of 30 seconds never pages; a real outage pages five minutes late. Guidelines:

A Prometheus restart does not reset a pending for clock: Prometheus stores the clocks in a series called ALERTS_FOR_STATE and restores them on startup, as long as it was down for less than --rules.alert.for-outage-tolerance (a command-line flag, 1h by default).

Symptoms, not causes

Chapter 0 made the argument (0.21); this is what it looks like in a rule file.

# cause-based: fires often, means little on its own
- alert: NodeHighCPU
  expr: 100 * (1 - avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m]))) > 80
  for: 5m
  labels: { severity: warning }

# symptom-based: fires when users are affected, whatever the cause
- alert: OrdersErrorBudgetBurn
  expr: job:slo_errors_per_request:ratio_rate1h > (14.4 * 0.001) and job:slo_errors_per_request:ratio_rate5m > (14.4 * 0.001)
  labels: { severity: page }

The first rule is the CPU query from 27.8 (100 minus the idle percentage) with > 80 on the end: it returns every instance above 80% busy. The second reads two recorded error ratios; the numbers 14.4 and 0.001 are burn-rate arithmetic from 0.17, built step by step in 28.14. For now read it as "orders is failing fast enough to spend its error budget, both over the last hour and right now".

High CPU during a batch job hurts nobody; a database failover that errors 8% of checkouts may not move CPU at all. Page on symptoms (errors, latency, availability as users see them). Keep cause alerts, but as tickets or dashboard annotations at low severity - they help diagnosis, they must not wake anyone.

Alert fatigue (0.21) comes from: pages nobody acts on, duplicate alerts for one problem (fixed with grouping and inhibition in Alertmanager, 28.7), flapping alerts, thresholds copied from a blog post, and alerts without a runbook. Every page should be urgent, actionable, and about users. If an alert fires and the right response is "ignore it", delete the alert.

Dead targets: up == 0 and absent()

Two different failures, two different expressions:

- alert: TargetDown                    # the target exists but cannot be scraped
  expr: up == 0
  for: 2m

- alert: NodeExporterAbsent            # the target is gone from service discovery
  expr: absent(up{job="node"})
  for: 5m

up == 0 needs an up series to exist. When a target disappears from the config entirely (a relabel rule that drops it, a ServiceMonitor with the wrong selector - both 27.4 - or a job someone renamed), there is no up series at all, the expression returns nothing, and the alert is silent exactly when monitoring is broken. absent() (27.15) turns "nothing" into one series with value 1, so it fires on exactly that. Write one absent(up{job="..."}) per job you cannot live without, and the same for key metrics: absent(http_server_requests_seconds_count{job="orders"}) catches the day an upgrade renames the metric.

There is also absent_over_time(x[10m]): it returns 1 if x had no samples at all in the last 10 minutes - handy for series that only appear now and then.

Disk will fill, and other predictions

- alert: DiskWillFillIn4Hours
  expr: |
    predict_linear(node_filesystem_avail_bytes{fstype!~"tmpfs|overlay"}[6h], 4 * 3600) < 0
      and
    node_filesystem_avail_bytes{fstype!~"tmpfs|overlay"} / node_filesystem_size_bytes < 0.2
  for: 30m
  labels:
    severity: ticket
  annotations:
    summary: "{{ $labels.mountpoint }} on {{ $labels.instance }} will be full within 4 hours"

Part by part: predict_linear(x[6h], 4 * 3600) (27.15) draws a straight line through the last 6 hours of free space and says where it will be in 4 hours (4 x 3600 seconds); < 0 keeps the disks predicted to run out. The fstype!~"tmpfs|overlay" matcher leaves out memory-backed and container filesystems. expr: | is YAML for "the next indented lines are one string", so a long expression can span lines.

The and clause keeps a left-hand series only if a series with the same labels also exists on the right. Here the right side is "less than 20% free", so the alert stays quiet on a disk that is 90% empty but briefly filling fast (a big download). and, or and unless match on labels like the arithmetic operators (27.15); both sides here have the same labels, so it just works.

Annotations that help at 3 a.m.

annotations:
  summary: "orders 5xx ratio is {{ $value | humanizePercentage }}"
  description: "{{ $labels.job }} has served more than 1.44% errors for 1 hour. Dashboard: https://dashboards.lab/d/orders"
  runbook_url: "https://wiki.lab/runbooks/orders-errors"

{{ $value }} is the value of the result series; for a comparison filter like x > 0.05 that is the left-hand side, x. The | passes it through humanizePercentage, which turns 0.0642 into 6.42%. A link to the dashboard and the runbook turns a page into a starting point.

Loading and checking

Rule files are listed under rule_files in prometheus.yml; this box uses a glob (a * pattern), so any /etc/prometheus/rules/*.yml is picked up on reload:

$ promtool check rules /etc/prometheus/rules/targets.yml
Checking /etc/prometheus/rules/targets.yml
  SUCCESS: 2 rules found

$ sudo systemctl reload prometheus
$ curl -s localhost:9090/api/v1/rules | jq -c '.data.groups[].rules[] | [.name, .state, .health]'
["NodeHighCPU","inactive","ok"]
["OrdersHighErrorRate","inactive","ok"]
["TargetDown","inactive","ok"]
["NodeExporterAbsent","inactive","ok"]

promtool check rules parses the file and every expression in it without running anything: SUCCESS: 2 rules found means the YAML and the PromQL are valid. The reload makes Prometheus read the files again. /api/v1/rules lists every loaded rule; the jq program prints, per rule, its name, its state (inactive, pending or firing - of the rule as a whole) and its health (ok, or err if the last evaluation failed).

"inactive, ok" is the healthy state of an alert that is not firing. It is also the state of an alert whose expression can never return anything - a typo in a label value, a regex that matches no series. Prometheus cannot tell a healthy system from a broken rule. Only a test can (next lesson, 28.4), and the second incident of this chapter (28.23) is exactly that.

What you can now do

Why it helps

You will write and review alert rules constantly, and the bugs are silent. A for: 1h on an error-rate alert means a 45-minute outage never pages. A {{ $value }} in a label creates a new alert every evaluation. An up == 0 alert says nothing when the target vanished from discovery; only absent(up{job="node"}) catches that. And "inactive, ok" in the rules API looks identical for a healthy system and for a rule that can never match.

When an incident review asks "did the alert fire?", max_over_time(ALERTS{alertname="X"}[3h]) answers it, including whether it was stuck pending. Symptoms versus causes is the argument you will make when a team wants to page on CPU. And a disk alert combining predict_linear with a percentage threshold is one you can copy into any environment.

Commands in this lesson

curl promtool systemctl

FAQ

What does the for field do exactly?

It is how long the expression must return a series continuously before the alert fires. On the first true evaluation the alert becomes pending with activeAt set; it fires at the first evaluation where now minus activeAt reaches for. A single false evaluation resolves it and resets the timer. Pending alerts stay in Prometheus; only firing ones are sent to Alertmanager.

Why shouldn't I put $value in a label?

Labels identify an alert instance. If a label contains the current value, it changes at almost every evaluation, so Prometheus sees a new alert each time: the old one resolves, a new one starts pending, and for never completes. Put {{ $value }} in annotations, which are free text, for example summary: "5xx ratio is {{ $value | humanizePercentage }}", and keep labels static like severity: page.

Why is my alert inactive when the system is clearly broken?

Either the condition is not met, or the expression can never return anything: a label value typo, an anchored regex like status=~"5xx", a metric that was renamed, a missing target so up == 0 has nothing to compare. Prometheus reports such a rule as inactive with health ok. Run the expression manually, check whether its selectors return series, and cover it with a promtool test.

What is keep_firing_for?

A rule option, available since Prometheus 2.42, that keeps an alert firing for a given time after its condition clears. It prevents resolve and refire storms from a noisy condition that flickers around the threshold, and gives a more stable notification. It does not help with conditions that are false before for elapses; for those, smooth the expression with a longer window or avg_over_time.

How do I find out afterwards whether an alert fired?

Prometheus writes a series called ALERTS with labels alertname and alertstate (pending or firing) for every active alert. max_over_time(ALERTS{alertname="TargetDown", alertstate="firing"}[3h]) returns 1 if it fired in the last three hours. Querying without the state filter also shows whether it was stuck pending. Alertmanager and the receiver's logs show whether a notification was actually sent.

In an interview Mid

Explain the lifecycle of a Prometheus alert.

An alerting rule is a PromQL expression evaluated every evaluation_interval; every series it returns is one alert instance (that is why up == 0 - a filter - returns exactly the dead targets). An empty result means all is fine.

Per instance:

You can watch it in /api/v1/alerts, and afterwards in the ALERTS series (alertstate pending or firing): max_over_time(ALERTS{alertname="X"}[3h]) tells you whether it ever fired.

Details that matter: labels must be static (a label with {{ $value }} creates a new alert each evaluation and for never completes); annotations are templates; "inactive, ok" in /api/v1/rules is also what a rule that can never match looks like.

Also asked: How do you detect that a service stopped being monitored at all? · What is the difference between symptom-based and cause-based alerting? · How do you choose the for duration of an alert?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.