An alert that has never fired is an untested alert
In 28.2 you broke a real exporter to see TargetDown fire. You cannot do that for every alert: you cannot wait for an outage to find out whether your page works, and you cannot create one on demand in production. What you want is what developers have for code - a unit test: a small, automatic check that feeds known input to one piece and compares the output with what you expect.
What you need to know already: alerting rules, for, pending and firing (28.1); counters and rate() (27.2, 27.8); recording rules (27.27); missing samples and staleness markers (27.2); CI pipelines (25.1); cat > file <<'EOF' heredocs (6.18).
The words you need first
- Synthetic series - fake data you write by hand, in place of real scrapes.
- Simulated clock - the test does not wait in real time; it jumps from minute to minute and evaluates the rules at each step, so an hour of alert behaviour takes a fraction of a second.
- Positive case / negative case - a check that the alert does fire when it should, and a check that it does not fire when it should not.
promtool test rules evaluates your real rule files against synthetic series you write, on a simulated clock, and checks which alerts fire when. It needs no running Prometheus.
# /home/learner/oncall-lab/labs/4d-observability/targets_test.yml
rule_files:
- /etc/prometheus/rules/targets.yml
evaluation_interval: 1m
tests:
- interval: 1m
input_series:
- series: 'up{job="node", instance="localhost:9100"}'
values: '1 1 1 0 0 0 0 0 0 0'
alert_rule_test:
- eval_time: 4m
alertname: TargetDown
exp_alerts: []
- eval_time: 6m
alertname: TargetDown
exp_alerts:
- exp_labels:
severity: warning
job: node
instance: localhost:9100
exp_annotations:
summary: "node target localhost:9100 is down"
$ promtool test rules targets_test.yml
Unit Testing: targets_test.yml
SUCCESS
The output names the test file and says SUCCESS: every expectation in it held.
Reading a test file
One field at a time:
rule_files- the rules under test. Relative paths are relative to the test file.evaluation_interval- how often rules run in the simulation (default 1m).tests- a list; each entry is an independent little world with its own data.interval- the spacing of the input samples (default: the evaluation interval).input_series- one entry per series: a label set written like a selector (up{job="node", instance="localhost:9100"}) and a string of values, one perinterval, starting at time 0.alert_rule_test- a list of checks. Each says: ateval_time(a moment on the simulated clock), which firing alerts namedalertnameexist.exp_alerts("expected alerts") lists them;exp_alerts: []asserts that none fire.promql_expr_test- ateval_time, what an arbitrary expression returns (great for testing recording rules; an example is further down).
So the file above says: the target is up for three minutes, then down. At minute 4 no TargetDown may be firing; at minute 6 exactly one, with these labels and this summary.
The values notation
Writing sixty numbers by hand would be painful, so promtool has a shorthand:
'1 1 1 0 0' five samples
'0+10x5' 0 10 20 30 40 50 (start + increment x times)
'100-5x3' 100 95 90 85
'1x4' 1 1 1 1 1 (repeat)
'0+60x60 3600+0x30' a counter growing 1/s for an hour, then flat
'_' a missing sample (the scrape did not happen)
'_x5' five missing samples
'stale' a staleness marker (the target went away)
Read 0+10x5 as "start at 0, add 10, five times": that gives six values. You can chain pieces with spaces, as in the fifth line.
Counters are written as their running totals, the way Prometheus sees them: 0+60x60 at interval: 1m is a counter that increases by 60 per minute, so rate(x[5m]) over it is 1 per second. To simulate an error ratio, write the errors and the total as two series with the right slopes:
- series: 'http_server_requests_seconds_count{job="orders", status="500", uri="/api/checkout"}'
values: '0+6x120' # 6 errors per minute
- series: 'http_server_requests_seconds_count{job="orders", status="200", uri="/api/checkout"}'
values: '0+54x120' # 54 successes per minute -> 10% errors
6 errors out of 60 requests a minute is 10%.
What the comparison checks - exactly
promtool compares the complete set of firing alerts for that name at that time, with all their labels and all their annotations:
exp_labelsare every label of the alert exceptalertname: the ones from the result series and the ones the rule adds (severity). Forget one and the test fails.exp_annotationsmust match the rendered annotations exactly (with the templates already filled in). If the rule has annotations and the test omitsexp_annotations, the expectation is "no annotations" and the test fails. This catches everybody once.- Pending alerts are not compared. At
eval_time: 4min the example the alert is pending (down since minute 3,for: 2m), soexp_alerts: []passes.
A failure prints both sides:
Unit Testing: targets_test.yml
FAILED:
alertname: TargetDown, time: 6m,
exp:[
0:
Labels:{alertname="TargetDown", instance="localhost:9100", job="node", severity="page"}
Annotations:{summary="node target localhost:9100 is down"}
],
got:[
0:
Labels:{alertname="TargetDown", instance="localhost:9100", job="node", severity="warning"}
Annotations:{summary="node target localhost:9100 is down"}
]
How to read it: alertname: TargetDown, time: 6m says which check failed. exp is what the test expected, got is what the rules really produced; the 0: is the first alert in each list. Compare them line by line: here the only difference is severity="page" against severity="warning" - the test expected the wrong severity.
Working out when an alert fires
Down from minute 3 (the fourth sample; the first is minute 0), for: 2m, evaluation every minute:
minute 3 up == 0 true -> pending, activeAt = 3m
minute 4 true -> 4 - 3 = 1m < 2m pending
minute 5 true -> 5 - 3 = 2m >= 2m FIRING
So it fires at 5m, and a test at eval_time: 5m expects it. The first true evaluation starts the clock; firing happens at the first evaluation where the elapsed time reaches for. Off-by-one errors in this arithmetic are the most common reason a test fails on the first try - that is the test doing its job.
Testing recording rules and edge cases
A recording rule has no alert to check, so you check its value with promql_expr_test:
promql_expr_test:
- expr: job:slo_errors_per_request:ratio_rate5m
eval_time: 30m
exp_samples:
- labels: 'job:slo_errors_per_request:ratio_rate5m{job="orders"}'
value: 0.1
exp_samples lists the series you expect (name and labels written as a selector) and each one's value. With the 6-errors-in-60 input above, the recorded ratio at minute 30 is 0.1.
Edge cases worth a test each, because they are where rules break:
- the gap:
values: '1 1 _ _ _ 1'- does a missing scrape makeup == 0fire? (No: missing is not zero. That is whatabsentis for.) - the reset: a counter that drops to 0 halfway (the app restarted, 27.8) - does the rate alert survive a deploy?
- no traffic:
0x60on both sides of a ratio - 0/0 is NaN ("not a number"), andNaN > xis false, so the error-ratio alert does not fire at night. Is that what you want? - the negative case: always assert
exp_alerts: []at a time just before it should fire. A test that only checks "it fires" passes for an alert that fires all the time.
In CI
promtool check rules and promtool test rules are fast, give the same result every time, and need no Prometheus. They belong in the pipeline (25.1) of whatever repository holds your rules, next to promtool check config, so a broken alert cannot be merged.
In Kubernetes (28.24) the same rules live inside PrometheusRule objects; you extract the rule part and run the same tests on it.
What you can now do
- Write a promtool test with synthetic series, positive and negative cases, and the exact labels and annotations.
- Work out the minute an alert with
forstarts firing, and prove it. - Read a FAILED diff and tell whether the rule or the test is wrong.