OnCallReady

Lesson 28.7 · Observability II: Alerting, Alertmanager & SLO alerts · 17 min read

Alertmanager: routing, grouping, inhibition, silences

In plain words

Imagine the school office when many alarms go off. The secretary does not phone every parent for every ring. She sorts: fire alarms go to the head teacher, broken taps to the caretaker. If twenty classrooms report the same power cut, she makes one call listing all twenty. If the whole building has no power, she does not also report every dark classroom. And if someone said "we are testing the bell from 2 to 4", she ignores the bell until then.

That is Alertmanager. The routing tree sends each alert to a receiver based on labels like severity="page". group_by bundles alerts into one notification, with group_wait, group_interval and repeat_interval as timers. Inhibition mutes warnings while a related page fires. Silences, created with amtool silence add, mute matching alerts for planned work.

What Alertmanager does

Picture twenty pods of one service failing at once: Prometheus fires twenty TargetDown alerts. Without help, the on-call phone rings twenty times, and the team chat gets the same twenty. Someone has to decide who hears about an alert, how often, and bundled with what - and that is not Prometheus's job.

What you need to know already: alerts, labels, pending and firing (28.1); label matchers =, !=, =~, !~ (27.8); external_labels in prometheus.yml (27.2); page vs ticket and alert fatigue (0.21); journalctl -u (2.30); systemctl reload and SIGHUP (2.21).

The words you need first

Prometheus decides that an alert fires. Alertmanager then:

  1. routes each alert down a tree to a receiver (pager, chat, email, webhook);
  2. groups alerts that belong together into one notification;
  3. throttles repeats (repeat_interval);
  4. inhibits alerts made redundant by a bigger one;
  5. applies silences people create for planned work.

Each of these gets its own section below.

On this box it is prometheus-alertmanager.service, listening on port 9093, configured in /etc/prometheus/alertmanager.yml. Prometheus knows where to send alerts because of the alerting.alertmanagers block in prometheus.yml (you saw it on the tour in 27.3).

The config on this box

global:
  resolve_timeout: 5m

route:
  receiver: team-chat
  group_by: ['alertname', 'job']
  group_wait: 30s
  group_interval: 5m
  repeat_interval: 4h
  routes:
    - matchers:
        - severity="page"
      receiver: oncall-pager

receivers:
  - name: team-chat
    webhook_configs:
      - url: http://127.0.0.1:5001/chat
  - name: oncall-pager
    webhook_configs:
      - url: http://127.0.0.1:5001/page

inhibit_rules:
  - source_matchers:
      - severity="page"
    target_matchers:
      - severity="warning"
    equal: ['instance']

A quick map before the details: route is the top of the routing tree (default receiver team-chat, with one child route for pages); the four group_* / repeat_* lines are the grouping timers; receivers defines the two destinations; inhibit_rules has one inhibition. resolve_timeout is how long Alertmanager keeps an alert it stopped hearing about before it treats it as resolved.

Both receivers are webhooks into pager-webhook, a lab stand-in for a paging service and a chat tool (simulator) that logs every notification it receives:

$ journalctl -u pager-webhook -n 2 --no-pager
... pager-webhook[17389]: CHAT receiver=team-chat status=firing group={alertname="TargetDown", job="node"} alerts=1
... pager-webhook[17389]:   [FIRING] TargetDown instance=localhost:9100 job=node monitor=oncall-lab severity=warning summary="node target localhost:9100 is down"

-n 2 shows the last two lines, --no-pager prints them straight to the terminal. The first line is one notification: it went to team-chat, it is about firing alerts, the group is "alertname TargetDown, job node", and it carries one alert. The indented line is that alert with all its labels and its summary.

(monitor="oncall-lab" comes from Prometheus's external_labels (27.2), added to every alert it sends. With several Prometheus servers, that is how you tell which one an alert came from.)

The routing tree

Every alert enters at the root route. Its child routes are tried in order; the first whose matchers all match the alert's labels takes the alert, and the search stops - unless that child has continue: true, in which case the next siblings are tried too. An alert no child matches stays at the parent (the default receiver). Settings a child does not give (group_by, the timers) are inherited from its parent.

amtool is Alertmanager's command-line tool. amtool config routes reads a config file and answers routing questions without touching the running server; --config.file= says which file:

$ amtool config routes show --config.file=/etc/prometheus/alertmanager.yml
Routing tree:
└── default-route  receiver: team-chat
    └── {severity="page"}  receiver: oncall-pager

$ amtool config routes test --config.file=/etc/prometheus/alertmanager.yml severity=page alertname=OrdersErrorBudgetBurn
oncall-pager
$ amtool config routes test --config.file=/etc/prometheus/alertmanager.yml severity=warning
team-chat

routes show draws the tree: the default route, and under it the page route with its matcher. routes test takes the labels of an imaginary alert as name=value arguments and prints the receiver it would land on. That is how you check a routing change before it goes live.

Matchers use the same syntax as PromQL label matchers: severity="page", team=~"orders|payments", env!="dev". (The older match: / match_re: maps still work and are deprecated - you will see them in old configs.)

Grouping and the three timers

group_by: ['alertname', 'job'] means all alerts with the same alertname and the same job go into one notification - a group. Twenty pods of the same service failing their health check is one message listing twenty, not twenty pages.

Three timers decide when a group's notifications go out:

group_by: ['...'] (the literal three dots) disables grouping: one notification per alert. Almost never what you want for pages.

Inhibition

An inhibit rule mutes target alerts while a source alert fires, if the labels listed in equal have the same values on both. The rule on this box: while any severity="page" alert fires for an instance, the severity="warning" alerts for the same instance are muted. The classic uses:

equal is the part people get wrong. Leave it out and one page anywhere mutes every warning everywhere. Put a label in it that the source alert does not have, and the inhibition never matches (both sides must have the same value, and "missing" only equals "missing").

Silences

A silence mutes every alert matching its matchers until it expires. Silences are created by people at run time (the web UI, the API, or amtool), not in the config file - they are for planned work like an upgrade:

$ amtool --alertmanager.url=http://localhost:9093 silence add alertname=TargetDown instance=localhost:9100 -d 2h -c 'kernel upgrade, OPS-1234'
e534d465-1d58-21d0-270c-cd4c32cbcc4c

$ amtool --alertmanager.url=http://localhost:9093 silence query
ID                                    Matchers                                          Ends At                  Created By  Comment
e534d465-1d58-21d0-270c-cd4c32cbcc4c  alertname="TargetDown" instance="localhost:9100"  2026-09-22 22:02:33 UTC  learner      kernel upgrade, OPS-1234

$ amtool --alertmanager.url=http://localhost:9093 alert query
Alertname  Starts At  Summary  State
$ amtool --alertmanager.url=http://localhost:9093 alert query -s
Alertname   Starts At                Summary                             State
TargetDown  2026-09-22 20:01:30 UTC  node target localhost:9100 is down  suppressed

These commands talk to the running Alertmanager, so they need --alertmanager.url= (where it listens). Piece by piece:

The rules for silences: always a comment with a ticket or reason, always the narrowest matchers that cover the work (not alertname=~".+", which matches every alert), always a duration, and expire it (amtool silence expire <id>) when you are done early. A forgotten broad silence is how real outages go unpaged.

Typing the URL every time gets old. Put it in ~/.config/amtool/config.yml once:

alertmanager.url: http://localhost:9093

Checking and reloading

$ amtool check-config /etc/prometheus/alertmanager.yml
Checking '/etc/prometheus/alertmanager.yml'  SUCCESS
Found:
 - global config
 - route
 - 1 inhibit rules
 - 2 receivers
 - 0 templates

amtool check-config validates a file without loading it: SUCCESS, then a count of what it found. (templates are optional files that customise the text of notifications; this box has none.)

A route pointing at a receiver name that is not in receivers: is the classic mistake. amtool prints FAILED: with the receiver's name and the route that uses it, then amtool: error: failed to validate 1 file(s) and exit code 1 - so it works as a gate in a CI pipeline.

sudo systemctl reload prometheus-alertmanager sends SIGHUP (2.21), which makes it re-read the file. Like Prometheus, a broken file is rejected and the old config keeps running; the journal says msg="Loading configuration file failed".

High availability, in one paragraph

If the one Alertmanager is down, nobody gets paged. So in production you run two or three and point every Prometheus at all of them. They talk to each other (a protocol called gossip, on port 9094) to share silences and to deduplicate notifications, so a page goes out once even though each Alertmanager received the alert. Do not put a load balancer (0.8, 9.23) in front of them for Prometheus to send to; the deduplication needs each one to see every alert.

Dashboard or alert?

If a human must act now, it is a page. If a human must act this week, it is a ticket. If it helps someone understand a problem they already know about, it belongs on a dashboard. CPU at 85%, a queue slightly longer than usual, GC pauses creeping up - dashboard. Checkout failing for users - page.

Later (Ch 29): you build those dashboards yourself, and add logs and traces for the "why" once a page has told you "what".

What you can now do

Why it helps

When on-call complains about twenty pages for one incident, the fix is group_by. When nobody got paged during an outage, you check with amtool config routes test whether the alert's labels landed on the chat receiver instead of the pager, and whether a forgotten broad silence swallowed it. When a node failure produces a storm of per-pod warnings, inhibition with the right equal labels is the answer, and without equal one page mutes every warning everywhere.

You will change Alertmanager config on a platform team, and amtool check-config catches the classic route to a receiver that does not exist before it ships. Understanding HA (all Prometheus servers send to all Alertmanagers, no load balancer) matters when you run it for real. Routing, grouping and inhibition are standard interview questions.

Commands in this lesson

journalctl amtool

FAQ

What is the difference between group_wait, group_interval and repeat_interval?

group_wait is how long Alertmanager waits after the first alert of a new group before sending, so alerts arriving together share one notification. group_interval is the minimum time before sending an update when new alerts join an already notified group. repeat_interval is how often an unchanged, still-firing group is re-sent. Typical values: 30s, 5m and a few hours.

What is the difference between a silence and an inhibition?

A silence is created at runtime by a person, through the UI, API or amtool, for planned work: it mutes alerts matching its matchers until it expires. An inhibition is a rule in the config: while a source alert fires, target alerts with the same values for the equal labels are muted automatically. Silences need a reason and an expiry; inhibitions encode known dependencies such as "node down makes pod alerts redundant".

Why did my inhibit rule mute everything, or nothing?

Without equal, any firing source alert mutes every target alert across all instances, clusters and teams. With a label in equal that the source alert lacks, both sides must have the same value, and "missing" only equals "missing", so it may never match. List the labels that tie cause to consequence, such as instance or cluster, and check that both alerts actually carry them.

How are routes matched?

Every alert enters at the root route. Child routes are checked in order and the first whose matchers match takes the alert; the search stops there unless that route has continue: true. Alerts matching no child stay at the parent. Children inherit settings they do not override. amtool config routes test with an alert's labels prints which receiver it would reach, which is the safe way to check a change.

How do I run Alertmanager in high availability?

Run two or three instances clustered with gossip (port 9094), and configure every Prometheus to send alerts to all of them. They share silences and deduplicate notifications, so a page is sent once. Do not put a load balancer between Prometheus and Alertmanager: deduplication depends on every instance receiving every alert. Prometheus itself is made HA by running two identical servers.

In an interview Mid

During an outage the alert fired in Prometheus but the on-call engineer was not paged. How do you investigate?

Follow the alert along its path, one hop at a time:

  1. Did it really fire? /api/v1/alerts or ALERTS{alertstate="firing"} - pending alerts never leave Prometheus.
  2. Did Alertmanager get it? amtool alert query - and -s, because silenced or inhibited alerts are hidden by default. State suppressed = received but muted.
  3. Was it muted? amtool silence query - a forgotten broad silence (alertname=~".+") is a classic. An inhibit rule without the right equal labels can mute everything.
  4. Where did it route? amtool config routes test severity=page job=orders prints the receiver. Routes are first-match: an earlier route may have swallowed it, or the severity label never said page.
  5. Timing: group_wait, group_interval and repeat_interval delay or batch notifications.
  6. Delivery: the receiver's own logs (the pager or webhook), and whether the last config reload failed - the journal says Loading configuration file failed while the old config keeps running.

Prevention: amtool check-config and routes test in CI, and two or three Alertmanagers that every Prometheus sends to.

Also asked: What does Alertmanager do? · What are grouping, inhibition and silences for in Alertmanager? · How would you design an Alertmanager routing tree for many teams?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.