OnCallReady

Lesson 27.4 · Observability I: Prometheus & PromQL · 17 min read

Service discovery and relabelling

In plain words

Imagine a pizza delivery driver. At first the boss gives a printed list of addresses every morning. But customers move, new ones appear, and the list is out of date by noon. So the boss connects the driver to the order system instead, which always knows the current addresses. Before each trip, the driver also applies a few rules to the list: "skip anyone who said no deliveries", "write just the street name on the ticket, not the flat number".

Service discovery is the order system: file_sd_configs, kubernetes_sd_configs and friends give Prometheus the current targets with hidden __meta_* labels. relabel_configs are the driver's rules, applied before the scrape: keep, drop, rewrite labels. metric_relabel_configs are applied after the scrape, to each series.

Targets that change by themselves

A static list of targets is fine for a box with three programs. On a platform, pods come and go every deploy and VMs are added by autoscaling. If a human has to edit prometheus.yml each time, a new service goes unmonitored until someone remembers - and a removed one stays in the list as a "down" target forever.

What you need to know already: targets, jobs, the job and instance labels, and reloading (27.2); regular expressions, including capture groups ( ) and $1 (7.3, 7.6); JSON (7.11); pods, labels and annotations in Kubernetes (15.14, 15.26).

Where targets come from

Service discovery (SD) is Prometheus asking something else for the list of targets instead of reading it from its own config. There is one mechanism per source: kubernetes_sd_configs (ask the Kubernetes API for pods), azure_sd_configs (ask Azure for VMs), consul_sd_configs, dns_sd_configs, and the most boring and most useful one, file_sd_configs: a JSON or YAML file that something else (Ansible, Terraform, a cron job) writes.

  - job_name: inventory
    file_sd_configs:
      - files:
          - /etc/prometheus/file_sd/inventory.json
[
  {
    "targets": ["inventory.lab:8000"],
    "labels": { "team": "inventory", "env": "lab" }
  }
]

The job says "read my targets from this file". The file is a JSON array; each object has a list of targets (host:port) and labels to put on all of them - the same shape as a static_configs entry.

Prometheus watches file_sd files: edit the JSON and the targets change within seconds, with no reload. Only prometheus.yml itself needs a reload. promtool check config warns when a listed file does not exist:

Checking /etc/prometheus/prometheus.yml
  WARNING: file "/etc/prometheus/file_sd/inventory.json" for file_sd in scrape job "inventory" does not exist
  SUCCESS: 2 rule files found

A WARNING, not a failure: the config is valid, the job simply has no targets until the file appears.

Discovered labels

Every discovered target starts with a set of labels, most of them hidden (their names start with __, two underscores):

"discoveredLabels": {
  "__address__": "inventory.lab:8000",
  "__metrics_path__": "/metrics",
  "__scheme__": "http",
  "__scrape_interval__": "15s",
  "__scrape_timeout__": "10s",
  "__meta_filepath": "/etc/prometheus/file_sd/inventory.json",
  "env": "lab",
  "job": "inventory",
  "team": "inventory"
}

__address__ is where to scrape; __scheme__ (http or https) and __metrics_path__ complete the URL: http://inventory.lab:8000/metrics. __meta_* labels are whatever the SD mechanism knows about the target - for file_sd only which file it came from; for Kubernetes the pod name, namespace, node, and every label and annotation of the pod. After the rewrite step below, everything starting with __ is thrown away, and if there is no instance label, instance is set to __address__.

You see them in /api/v1/targets as discoveredLabels (before the rewrite) and labels (after). When the rewrite does not do what you think, compare those two.

relabel_configs: rewrite targets before the scrape

Relabelling is Prometheus rewriting labels with rules you give it. relabel_configs (inside a job) is an ordered list of rules applied to each target's labels, once, before it is scraped. Each rule:

relabel_configs:
  # instance = host part of the address, without the port
  - source_labels: [__address__]
    regex: '(.+):\d+'
    target_label: instance
    replacement: '$1'

  # scrape only targets that opted in (Kubernetes annotation convention)
  - source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_scrape]
    action: keep
    regex: 'true'

  # never scrape the test environment
  - source_labels: [env]
    action: drop
    regex: 'test|dev'

  # copy every pod label into a real label
  - action: labelmap
    regex: '__meta_kubernetes_pod_label_(.+)'

  # remove a label nobody needs
  - action: labeldrop
    regex: 'pod_template_hash'

Take the first rule apart, because the next mission uses it: the source is __address__ = inventory.lab:8000. The regex (.+):\d+ means "one or more of anything, captured as group 1, then a colon, then digits" - it matches the whole string with group 1 = inventory.lab. The rule writes the replacement ($1, group 1) into target_label instance. Result: instance="inventory.lab", no port.

The actions you will actually use:

replace    (default) regex on source_labels -> write replacement to target_label
keep       drop the target unless the joined source labels match regex
drop       drop the target if they match
labelmap   rename labels whose NAME matches regex (replacement builds the new name)
labeldrop  delete labels whose NAME matches
labelkeep  delete every label whose NAME does not match
hashmod    target_label = hash(source) % modulus - sharding across Prometheus servers

(hashmod splits targets between several Prometheus servers so each scrapes a share; you will rarely write it yourself.)

Two traps with replace:

  1. The regex must match the whole value. regex: ':\d+' does not match inventory.lab:8000 (the host part is not covered), so nothing happens and nothing tells you.
  2. $1_suffix is not "group 1 then _suffix". Go reads it as a group named 1_suffix, which does not exist, so you get an empty string. Write ${1}_suffix.

metric_relabel_configs: rewrite series after the scrape

Same syntax, different moment: metric_relabel_configs is applied to every scraped series before it is stored. The metric name is a label too, called __name__. This is the tool for protecting Prometheus from an application that exposes something it should not:

metric_relabel_configs:
  # drop a whole metric
  - source_labels: [__name__]
    regex: 'go_gc_duration_seconds.*'
    action: drop

  # drop the series of one metric whose path label has an id in it
  - source_labels: [__name__, path]
    regex: 'http_requests_total;/api/items/[0-9]+'
    action: drop

The second rule joins two labels with ;, so for a series http_requests_total{path="/api/items/1417"} the string it tests is http_requests_total;/api/items/1417, which matches - dropped. A series with path="/healthz" does not match and is kept.

It is a stopgap. The fix is in the application (use the route template, /api/items/{id}, as the label, which is what Spring does with uri). And it costs CPU: Prometheus still parses every line before throwing it away.

Do not try to fix too many series (27.1's cardinality) with labeldrop on the offending label. Six hundred series that differ only by path become six hundred series with the same labels; Prometheus keeps one sample per timestamp and discards the rest as duplicates, and your counter is now a random one of the six hundred.

Target labels vs exposed labels: honor_labels

If a target exposes a label that Prometheus also sets (job, instance, or one from labels:), the target's copy is renamed exported_job (and so on) and Prometheus's wins. honor_labels: true reverses that. It exists for the Pushgateway (27.2) and for one Prometheus scraping another, where the exposed job is the real one.

Kubernetes, at concept level

In a cluster you almost never write kubernetes_sd_configs by hand. The Prometheus Operator (an operator is a program running in the cluster that installs and manages an app for you, driven by its own resource types) turns a ServiceMonitor or PodMonitor object - "scrape the pods behind Services labelled app: orders" - into exactly the relabelling above:

apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
  name: orders
spec:
  selector:
    matchLabels: { app: orders }
  endpoints:
    - port: http
      path: /actuator/prometheus
      interval: 15s

selector picks the Services by label (16.1), endpoints says which named port and path to scrape and how often. It is still relabelling underneath. When a ServiceMonitor "does nothing", the debugging is the same as here: look at /api/v1/targets, check whether the target is in droppedTargets (targets a keep/drop rule removed), and compare discovered labels with the selector.

What you can now do

Why it helps

In a cluster, targets come from discovery, and "my ServiceMonitor does nothing" is a weekly ticket on a platform team. Debugging it is always the same: find the target in droppedTargets, compare discoveredLabels with the final labels, and see which relabel rule dropped it. Knowing that regexes are fully anchored and that $1_suffix is read as a group named 1_suffix explains most silent failures.

metric_relabel_configs is your emergency brake when a team deploys a metric with ids in the path and Prometheus memory climbs: drop the series now, fix the instrumentation later. And knowing why labeldrop on the offending label corrupts the data, instead of fixing cardinality, saves you from making things worse at 3am. file_sd with a file written by Terraform or Ansible is the simplest way to monitor VMs.

FAQ

What is the difference between relabel_configs and metric_relabel_configs?

relabel_configs runs on each target's labels before the scrape: it decides whether to scrape the target and what its instance, job and other target labels will be, using __meta_* labels from discovery. metric_relabel_configs runs after the scrape on every scraped series, before storage: it can drop or rewrite metrics. Same syntax, different moment and different cost.

Why does my relabel regex not match?

Prometheus anchors relabel regexes at both ends: regex: foo means ^(?:foo)$. So ':\d+' does not match inventory.lab:8000; you need '(.+):\d+'. Source labels are joined with ; before matching, which matters when you list several. Also, in the replacement, $1_suffix is read as a group named 1_suffix; write ${1}_suffix. None of these produce an error.

Do I need to reload Prometheus when a file_sd file changes?

No. Prometheus watches file_sd files and picks up changes within seconds. Only changes to prometheus.yml itself, including adding a new file_sd job, need a reload. promtool check config warns if a referenced file does not exist. That makes file_sd a good integration point: Terraform, Ansible or a cron job writes the JSON, and Prometheus follows.

What does honor_labels do?

When a scraped series has a label that Prometheus also sets, such as job or instance, by default Prometheus keeps its own and renames the target's to exported_job. With honor_labels: true the target's labels win. That is meant for the Pushgateway and federation, where the exposed job is the real origin. For normal targets leave it off, so targets cannot overwrite your labels.

How does a ServiceMonitor relate to relabelling?

A ServiceMonitor is a Prometheus Operator object that says "scrape the endpoints of Services matching these labels, on this port and path". The operator turns it into a scrape job with kubernetes_sd_configs and relabel rules that keep matching endpoints and set labels like namespace, pod and service. When it does not work, debug it like any relabelling: check the targets API, droppedTargets, and whether the Service labels and port name match.

In an interview Mid

What is service discovery in Prometheus, and what is relabelling for?

Service discovery = Prometheus asks something else for its list of targets instead of a static list: kubernetes_sd_configs (pods), azure_sd_configs (VMs), dns_sd_configs, or file_sd_configs (a JSON/YAML file another tool writes; Prometheus watches it, no reload needed). New pods and VMs are monitored without anyone editing prometheus.yml, and removed ones disappear.

Each discovered target comes with hidden labels: __address__, __metrics_path__, __scheme__, and __meta_* (for Kubernetes: pod, namespace, every label and annotation). Relabelling turns those into the target you want:

Debug by comparing discoveredLabels with labels in /api/v1/targets (and droppedTargets). A ServiceMonitor is the same relabelling underneath.

Also asked: A team deployed a new version and Prometheus memory is climbing fast. What do you do? · What is the difference between relabel_configs and metric_relabel_configs? · A ServiceMonitor seems to do nothing. How do you debug it?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.