OnCallReady

Lesson 28.4 · Observability II: Alerting, Alertmanager & SLO alerts · 16 min read

promtool test rules: prove the alert fires

In plain words

Imagine testing a fire alarm. You do not set the school on fire; you hold a little smoke spray under it and check that it beeps, and also check it does not beep when someone just makes toast. You write down exactly what should happen: "after 30 seconds of smoke, beep; before that, silent".

promtool test rules is the smoke spray. You write synthetic series such as up{job="node"} with values '1 1 1 0 0 0', run your real rule file on a pretend clock, and assert at eval_time: 6m that TargetDown fires with exactly these labels and annotations, and at 4m that it does not. No Prometheus server and no real outage needed, so it runs in CI on every change to the rules.

An alert that has never fired is an untested alert

In 28.2 you broke a real exporter to see TargetDown fire. You cannot do that for every alert: you cannot wait for an outage to find out whether your page works, and you cannot create one on demand in production. What you want is what developers have for code - a unit test: a small, automatic check that feeds known input to one piece and compares the output with what you expect.

What you need to know already: alerting rules, for, pending and firing (28.1); counters and rate() (27.2, 27.8); recording rules (27.27); missing samples and staleness markers (27.2); CI pipelines (25.1); cat > file <<'EOF' heredocs (6.18).

The words you need first

promtool test rules evaluates your real rule files against synthetic series you write, on a simulated clock, and checks which alerts fire when. It needs no running Prometheus.

# /home/learner/oncall-lab/labs/4d-observability/targets_test.yml
rule_files:
  - /etc/prometheus/rules/targets.yml

evaluation_interval: 1m

tests:
  - interval: 1m
    input_series:
      - series: 'up{job="node", instance="localhost:9100"}'
        values: '1 1 1 0 0 0 0 0 0 0'
    alert_rule_test:
      - eval_time: 4m
        alertname: TargetDown
        exp_alerts: []
      - eval_time: 6m
        alertname: TargetDown
        exp_alerts:
          - exp_labels:
              severity: warning
              job: node
              instance: localhost:9100
            exp_annotations:
              summary: "node target localhost:9100 is down"
$ promtool test rules targets_test.yml
Unit Testing:  targets_test.yml
  SUCCESS

The output names the test file and says SUCCESS: every expectation in it held.

Reading a test file

One field at a time:

So the file above says: the target is up for three minutes, then down. At minute 4 no TargetDown may be firing; at minute 6 exactly one, with these labels and this summary.

The values notation

Writing sixty numbers by hand would be painful, so promtool has a shorthand:

'1 1 1 0 0'          five samples
'0+10x5'             0 10 20 30 40 50          (start + increment x times)
'100-5x3'            100 95 90 85
'1x4'                1 1 1 1 1                 (repeat)
'0+60x60 3600+0x30'  a counter growing 1/s for an hour, then flat
'_'                  a missing sample (the scrape did not happen)
'_x5'                five missing samples
'stale'              a staleness marker (the target went away)

Read 0+10x5 as "start at 0, add 10, five times": that gives six values. You can chain pieces with spaces, as in the fifth line.

Counters are written as their running totals, the way Prometheus sees them: 0+60x60 at interval: 1m is a counter that increases by 60 per minute, so rate(x[5m]) over it is 1 per second. To simulate an error ratio, write the errors and the total as two series with the right slopes:

- series: 'http_server_requests_seconds_count{job="orders", status="500", uri="/api/checkout"}'
  values: '0+6x120'       # 6 errors per minute
- series: 'http_server_requests_seconds_count{job="orders", status="200", uri="/api/checkout"}'
  values: '0+54x120'      # 54 successes per minute  ->  10% errors

6 errors out of 60 requests a minute is 10%.

What the comparison checks - exactly

promtool compares the complete set of firing alerts for that name at that time, with all their labels and all their annotations:

A failure prints both sides:

Unit Testing:  targets_test.yml
  FAILED:
    alertname: TargetDown, time: 6m,
        exp:[
            0:
              Labels:{alertname="TargetDown", instance="localhost:9100", job="node", severity="page"}
              Annotations:{summary="node target localhost:9100 is down"}
            ],
        got:[
            0:
              Labels:{alertname="TargetDown", instance="localhost:9100", job="node", severity="warning"}
              Annotations:{summary="node target localhost:9100 is down"}
            ]

How to read it: alertname: TargetDown, time: 6m says which check failed. exp is what the test expected, got is what the rules really produced; the 0: is the first alert in each list. Compare them line by line: here the only difference is severity="page" against severity="warning" - the test expected the wrong severity.

Working out when an alert fires

Down from minute 3 (the fourth sample; the first is minute 0), for: 2m, evaluation every minute:

minute 3   up == 0 true   -> pending, activeAt = 3m
minute 4   true           -> 4 - 3 = 1m  < 2m  pending
minute 5   true           -> 5 - 3 = 2m >= 2m  FIRING

So it fires at 5m, and a test at eval_time: 5m expects it. The first true evaluation starts the clock; firing happens at the first evaluation where the elapsed time reaches for. Off-by-one errors in this arithmetic are the most common reason a test fails on the first try - that is the test doing its job.

Testing recording rules and edge cases

A recording rule has no alert to check, so you check its value with promql_expr_test:

    promql_expr_test:
      - expr: job:slo_errors_per_request:ratio_rate5m
        eval_time: 30m
        exp_samples:
          - labels: 'job:slo_errors_per_request:ratio_rate5m{job="orders"}'
            value: 0.1

exp_samples lists the series you expect (name and labels written as a selector) and each one's value. With the 6-errors-in-60 input above, the recorded ratio at minute 30 is 0.1.

Edge cases worth a test each, because they are where rules break:

In CI

promtool check rules and promtool test rules are fast, give the same result every time, and need no Prometheus. They belong in the pipeline (25.1) of whatever repository holds your rules, next to promtool check config, so a broken alert cannot be merged.

In Kubernetes (28.24) the same rules live inside PrometheusRule objects; you extract the rule part and run the same tests on it.

What you can now do

Why it helps

An alert that has never fired is an untested alert, and the second incident of this chapter is a rule that could never fire. Tests catch the bugs you cannot see in review: the anchored regex, the off-by-one in for arithmetic, the label that the and does not match, the annotation template that renders wrong.

On a platform team the rules live in a repo, and promtool check rules plus promtool test rules in the pipeline is how you let many teams change alerts without breaking paging. Writing the edge-case tests (a missing scrape, a counter reset during a deploy, zero traffic at night giving NaN) forces you to decide the behaviour on purpose. Interviewers who ask "how do you know your alerts work?" are looking for exactly this answer.

Commands in this lesson

promtool

FAQ

What does the values notation 0+60x60 mean?

It is start, increment and count: 0+60x60 means 0, 60, 120, and so on, 61 values in total, one per interval. With interval: 1m that is a counter growing 60 per minute, so rate() over it gives 1 per second. 100-5x3 counts down, 1x4 repeats 1 five times, _ is a missing sample, and stale is a staleness marker as if the target disappeared.

Why does my test fail on annotations I didn't specify?

promtool compares the complete set of firing alerts for that name at that time, with all labels and all annotations. If the rule has annotations and the test omits exp_annotations, the expectation is "no annotations", which does not match. Include every annotation exactly as rendered, and every label except alertname, including the ones the rule adds such as severity.

Are pending alerts checked by alert_rule_test?

No. exp_alerts lists only firing alerts. An alert that is still pending at eval_time counts as not firing, so exp_alerts: [] passes during the for period. That is useful for negative tests just before the alert should fire, and a common reason tests pass when you expected them to fail: the alert was pending, not firing, at that time.

How do I test a recording rule?

Use promql_expr_test: give an expression, often the recorded series name, an eval_time and the expected samples with their labels and values. For example, with input series that produce 6 errors and 54 successes per minute, the recorded error ratio at 30m should be 0.1. This verifies the aggregation labels as well as the arithmetic, which matters because alerts match recorded series by labels.

Can I test PrometheusRule objects from Kubernetes?

Yes. The spec of a PrometheusRule is the same format as a rule file, so extract it with yq, a jq-like tool for YAML: yq '.spec' rule.yaml > rules.yml and run promtool check rules and promtool test rules against the extracted file in CI. The test files reference the extracted rule file in rule_files. This is a common pattern for repos holding kube-prometheus-stack rules.

In an interview Mid

How do you test Prometheus alerting rules?

With promtool test rules: it evaluates the real rule files against synthetic series you write, on a simulated clock, with no running Prometheus.

A test file lists rule_files, then tests with input_series (a label set plus values in the shorthand - '1 1 1 0 0', '0+60x60' for a counter growing 60 per interval, _ for a missing sample) and alert_rule_test entries: at eval_time, exactly these firing alerts with these exp_labels and exp_annotations. promql_expr_test checks a recording rule's value.

What makes the tests worth having:

Run promtool check rules and test rules in the CI pipeline of the rules repo, so a broken alert cannot be merged.

Also asked: An alert rule passed code review but never fired during a real outage. How do you prevent this class of bug? · Why should every alert test include a negative case? · How would you check alert rules automatically before they are merged?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.