OnCallReady

Lesson 27.27 · Observability I: Prometheus & PromQL · 13 min read

Recording rules

In plain words

Imagine your class keeps asking the teacher the same hard sum: "how many kids from all classes were late this week?" Every time, the teacher walks to every classroom and counts again. Smarter: a helper counts once every morning and writes the answer on the board. Anyone who asks just reads the board. The catch: the board only has numbers from the day the helper started; nobody wrote down last month.

A recording rule is that helper. Prometheus evaluates an expression like the error ratio per job and uri every 15 seconds and stores the result as a new series, named by convention job_uri:http_server_requests_errors:ratio_rate5m. Dashboards and alerts read the cheap stored series instead of recomputing, and the series has no history before the rule was loaded.

Computing it once instead of every time

The p99-by-endpoint query touches every bucket series of every instance. Put it on a dashboard that refreshes every 10 seconds, open it on five laptops, and Prometheus recomputes the same answer 30 times a minute. And when the next chapter writes alerts from the same error ratio at six different windows, you want that ratio written down once, the same way everywhere.

What you need to know already: sum by, rate and error ratios (27.8, 27.15); histogram_quantile (27.15); rule_files and evaluation_interval in prometheus.yml, promtool check config and reloading (27.2); YAML multi-line strings with | (11.31).

Recording rules

A rule is a query Prometheus runs by itself on a schedule (evaluation_interval, 15 s here). A recording rule stores each result as a new time series with a name you choose. Two reasons to write one:

  1. Cost. A dashboard that computes histogram_quantile(0.99, sum by (le, uri) (rate(...[5m]))) touches every bucket series on every refresh, for every viewer. A recording rule computes it once per 15 s, and the dashboard reads one cheap series.
  2. Building blocks. Alerts and SLOs (0.8) are built on other queries. Writing each ratio once as a recorded series keeps everything that uses it short and consistent.

(The other kind of rule, an alerting rule, sends an alert when its query returns something. Recording and alerting rules live in the same files; the next chapter is about the alerting kind.)

groups:
  - name: orders-recording
    rules:
      - record: job_uri:http_server_requests_seconds_count:rate5m
        expr: sum by (job, uri) (rate(http_server_requests_seconds_count[5m]))

      - record: job_uri:http_server_requests_errors:ratio_rate5m
        expr: |
          sum by (job, uri) (rate(http_server_requests_seconds_count{status=~"5.."}[5m]))
            /
          sum by (job, uri) (rate(http_server_requests_seconds_count[5m]))

      - record: job_uri:http_server_requests_seconds:p99_5m
        expr: histogram_quantile(0.99, sum by (job, uri, le) (rate(http_server_requests_seconds_bucket[5m])))

Each rule has two keys: record (the name of the new series) and expr (the PromQL). expr: | starts a YAML multi-line string, so a long query can be split over indented lines. These three are the request rate, the error ratio and the p99, each per job and uri.

The naming convention

level:metric:operations, from the Prometheus docs:

job_uri:http_server_requests_errors:ratio_rate5m tells you, without opening the file, that it is an error ratio over 5 minutes, one series per job and uri. The colons are legal in metric names and by convention used only by recording rules, so anyone reading a query knows it is not raw data.

Rule files and groups

Check before you load. promtool check rules <file> parses a rule file and every query in it:

$ promtool check rules /etc/prometheus/rules/recording.yml
Checking /etc/prometheus/rules/recording.yml
  SUCCESS: 3 rules found

A broken expression is reported with the file position and the parser's message:

$ promtool check rules /etc/prometheus/rules/recording.yml
Checking /etc/prometheus/rules/recording.yml
  FAILED:
/etc/prometheus/rules/recording.yml: 5:15: group "orders-recording", rule 1, "job_uri:http_server_requests_seconds_count:rate5m": could not parse expression: 1:55: parse error: unexpected end of input in function call, expected ")"

Reading it: file line 5, column 15; which group and which rule; then the PromQL parser's own message, with a position inside the query (column 55: a missing )).

And a typo in a field name is a YAML error, naming the Go type it expected:

/etc/prometheus/rules/recording.yml: yaml: unmarshal errors:
  line 5: field exprr not found in type rulefmt.RuleNode

promtool check config checks the rule files the config references, too - that is the one to run before systemctl reload prometheus.

Recording rules have no past

A recorded series starts existing at the first evaluation after the rule is loaded. Nothing is filled in for the past (no backfill). Right after a reload:

$ promtool query instant http://localhost:9090 'job_uri:http_server_requests_errors:ratio_rate5m'
# (nothing yet: the first evaluation has not happened)

Fifteen seconds later there is data, and [1h] over it has one hour of history only an hour from now. Consequences:

Is it running? The rules API

/api/v1/rules lists every loaded rule group and whether each rule's last evaluation worked:

# after the recording-rules mission loads recording.yml
curl -s localhost:9090/api/v1/rules | jq -r '.data.groups[] | .name as $g | .rules[] | [$g, .type, .name, .health] | @tsv'
orders-recording  recording  job_uri:http_server_requests_seconds_count:rate5m  ok
orders-recording  recording  job_uri:http_server_requests_errors:ratio_rate5m   ok
orders-recording  recording  job_uri:http_server_requests_seconds:p99_5m        ok
node              alerting   NodeHighCPU                                         ok
orders            alerting   OrdersHighErrorRate                                 ok

The jq program: for each group, remember its name as $g (as $g, 7.13), then print group, type, rule name and health per rule. Your three recording rules are ok; the two alerting rules came with the box.

health: err with a lastError means the expression failed at evaluation time (a many-to-many match that only happens with real data, for instance). A rule that parses but selects nothing is ok and records nothing - the same silent failure as a wrong regex in a query.

What you can now do

Why it helps

When dashboards time out or Prometheus CPU spikes every time the team opens the overview, the fix is recording rules for the expensive histogram and ratio queries. When you build the SLO burn-rate alerts in the next chapter, you need the error ratio at six window sizes, and writing each once as a recorded series is what keeps the alert rules readable and consistent.

The traps are practical: a new recorded series is empty right after deploy, so an alert on [1h] of it cannot fire for an hour, which confuses anyone testing it. A rule that parses but selects nothing is health: ok and records nothing. And the naming convention lets you read a colleague's query and know instantly what level and window a series is, which also shows up in interview answers about scaling Prometheus.

Commands in this lesson

promtool

FAQ

When should I create a recording rule?

When a query is expensive and used often, like a histogram_quantile over many bucket series on a dashboard many people open, or when an expression is a building block reused by several alerts or other rules, like error ratios at several windows for SLO alerts. Also for federation and remote write, where you send pre-aggregated series instead of raw ones. Do not record every query; each recorded series costs storage.

What does the name level:metric:operations mean?

It is the Prometheus naming convention for recorded series. level lists the labels the result is aggregated to (job_uri), metric is the source metric name with _total stripped after a rate, and operations lists what was done, newest last (ratio_rate5m, p99_5m). Colons appear only in recorded series, so anyone reading a query knows it is not raw data.

Why is my new recorded series empty?

A recording rule starts producing data at its first evaluation after being loaded, and there is no backfill. Right after a reload the series does not exist; after one evaluation interval it has one sample; [1h] over it has a full hour only an hour later. If it stays empty, check the rules API: health: ok with no data means the expression selects nothing, for example a wrong label or regex.

Can a rule use a series recorded by another rule?

Yes. Within a group, rules are evaluated sequentially at the same evaluation time, so a later rule sees the value an earlier rule recorded in the same cycle. Across groups, which run in parallel on their own intervals, a rule sees the other group's result from its latest evaluation, possibly one interval old. Put dependent rules in the same group, in order.

How do I check rule files before loading them?

promtool check rules file.yml parses each rule and reports expression errors with file position and parser message, and YAML field typos like field exprr not found in type rulefmt.RuleNode. promtool check config prometheus.yml also checks all rule files the config references, so run it before systemctl reload prometheus. After reloading, /api/v1/rules shows each rule's health and last error.

In an interview Mid

What are recording rules in Prometheus and why would you use them?

A recording rule is a query Prometheus evaluates on its own every evaluation_interval and stores as a new time series:

groups:
  - name: orders-recording
    rules:
      - record: job_uri:http_server_requests_errors:ratio_rate5m
        expr: |
          sum by (job, uri) (rate(http_server_requests_seconds_count{status=~"5.."}[5m]))
            / sum by (job, uri) (rate(http_server_requests_seconds_count[5m]))

Why:

  1. Cost - an expensive query (a p99 over every bucket series) is computed once per interval instead of on every dashboard refresh for every viewer.
  2. Building blocks - alerts and SLOs reuse one consistently defined ratio.

Naming: level:metric:operations - aggregated to job_uri, from that metric, ratio_rate5m done to it. Colons mark recorded series.

Operating them: promtool check rules (and check config) before systemctl reload prometheus; /api/v1/rules shows each rule's health. And the gotcha: a recorded series has no past - it starts at the first evaluation after loading, so anything using [1h] of it is unreliable for the first hour.

Also asked: You deployed an alert based on a recorded series and it does not fire during testing. Why? · How should recording rules be named, and why? · How do you check that a rule file is valid and that its rules are running?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.