What Alertmanager does
Picture twenty pods of one service failing at once: Prometheus fires twenty TargetDown alerts. Without help, the on-call phone rings twenty times, and the team chat gets the same twenty. Someone has to decide who hears about an alert, how often, and bundled with what - and that is not Prometheus's job.
What you need to know already: alerts, labels, pending and firing (28.1); label matchers =, !=, =~, !~ (27.8); external_labels in prometheus.yml (27.2); page vs ticket and alert fatigue (0.21); journalctl -u (2.30); systemctl reload and SIGHUP (2.21).
The words you need first
- Alertmanager - a separate program that receives firing alerts from one or more Prometheus servers and turns them into messages for people.
- Notification - one message Alertmanager sends (a page, a chat message, an email). One notification can carry many alerts.
- Receiver - a named destination for notifications: a pager service, a chat channel, email, or a webhook (an HTTP URL Alertmanager POSTs the alerts to as JSON). PagerDuty and Opsgenie are common paging services (they phone the on-call person); Slack and Teams are common chat tools.
Prometheus decides that an alert fires. Alertmanager then:
- routes each alert down a tree to a receiver (pager, chat, email, webhook);
- groups alerts that belong together into one notification;
- throttles repeats (
repeat_interval); - inhibits alerts made redundant by a bigger one;
- applies silences people create for planned work.
Each of these gets its own section below.
On this box it is prometheus-alertmanager.service, listening on port 9093, configured in /etc/prometheus/alertmanager.yml. Prometheus knows where to send alerts because of the alerting.alertmanagers block in prometheus.yml (you saw it on the tour in 27.3).
The config on this box
global:
resolve_timeout: 5m
route:
receiver: team-chat
group_by: ['alertname', 'job']
group_wait: 30s
group_interval: 5m
repeat_interval: 4h
routes:
- matchers:
- severity="page"
receiver: oncall-pager
receivers:
- name: team-chat
webhook_configs:
- url: http://127.0.0.1:5001/chat
- name: oncall-pager
webhook_configs:
- url: http://127.0.0.1:5001/page
inhibit_rules:
- source_matchers:
- severity="page"
target_matchers:
- severity="warning"
equal: ['instance']
A quick map before the details: route is the top of the routing tree (default receiver team-chat, with one child route for pages); the four group_* / repeat_* lines are the grouping timers; receivers defines the two destinations; inhibit_rules has one inhibition. resolve_timeout is how long Alertmanager keeps an alert it stopped hearing about before it treats it as resolved.
Both receivers are webhooks into pager-webhook, a lab stand-in for a paging service and a chat tool (simulator) that logs every notification it receives:
$ journalctl -u pager-webhook -n 2 --no-pager
... pager-webhook[17389]: CHAT receiver=team-chat status=firing group={alertname="TargetDown", job="node"} alerts=1
... pager-webhook[17389]: [FIRING] TargetDown instance=localhost:9100 job=node monitor=oncall-lab severity=warning summary="node target localhost:9100 is down"
-n 2 shows the last two lines, --no-pager prints them straight to the terminal. The first line is one notification: it went to team-chat, it is about firing alerts, the group is "alertname TargetDown, job node", and it carries one alert. The indented line is that alert with all its labels and its summary.
(monitor="oncall-lab" comes from Prometheus's external_labels (27.2), added to every alert it sends. With several Prometheus servers, that is how you tell which one an alert came from.)
The routing tree
Every alert enters at the root route. Its child routes are tried in order; the first whose matchers all match the alert's labels takes the alert, and the search stops - unless that child has continue: true, in which case the next siblings are tried too. An alert no child matches stays at the parent (the default receiver). Settings a child does not give (group_by, the timers) are inherited from its parent.
amtool is Alertmanager's command-line tool. amtool config routes reads a config file and answers routing questions without touching the running server; --config.file= says which file:
$ amtool config routes show --config.file=/etc/prometheus/alertmanager.yml
Routing tree:
└── default-route receiver: team-chat
└── {severity="page"} receiver: oncall-pager
$ amtool config routes test --config.file=/etc/prometheus/alertmanager.yml severity=page alertname=OrdersErrorBudgetBurn
oncall-pager
$ amtool config routes test --config.file=/etc/prometheus/alertmanager.yml severity=warning
team-chat
routes show draws the tree: the default route, and under it the page route with its matcher. routes test takes the labels of an imaginary alert as name=value arguments and prints the receiver it would land on. That is how you check a routing change before it goes live.
Matchers use the same syntax as PromQL label matchers: severity="page", team=~"orders|payments", env!="dev". (The older match: / match_re: maps still work and are deprecated - you will see them in old configs.)
Grouping and the three timers
group_by: ['alertname', 'job'] means all alerts with the same alertname and the same job go into one notification - a group. Twenty pods of the same service failing their health check is one message listing twenty, not twenty pages.
Three timers decide when a group's notifications go out:
group_wait(30s) - after the first alert of a new group arrives, wait this long before notifying, so the others arriving in the same moment make it into the first message.group_interval(5m) - after a notification, wait at least this long before sending an update about new alerts in the same group.repeat_interval(4h) - resend an unchanged, still-firing group this often. Too short is nagging; too long and a page acknowledged and forgotten stays forgotten.
group_by: ['...'] (the literal three dots) disables grouping: one notification per alert. Almost never what you want for pages.
Inhibition
An inhibit rule mutes target alerts while a source alert fires, if the labels listed in equal have the same values on both. The rule on this box: while any severity="page" alert fires for an instance, the severity="warning" alerts for the same instance are muted. The classic uses:
- the whole cluster/zone is down -> mute every per-service alert in it;
- a node is down -> mute the alerts for every pod on that node;
- a page fires -> mute its own lower-severity early warning.
equal is the part people get wrong. Leave it out and one page anywhere mutes every warning everywhere. Put a label in it that the source alert does not have, and the inhibition never matches (both sides must have the same value, and "missing" only equals "missing").
Silences
A silence mutes every alert matching its matchers until it expires. Silences are created by people at run time (the web UI, the API, or amtool), not in the config file - they are for planned work like an upgrade:
$ amtool --alertmanager.url=http://localhost:9093 silence add alertname=TargetDown instance=localhost:9100 -d 2h -c 'kernel upgrade, OPS-1234'
e534d465-1d58-21d0-270c-cd4c32cbcc4c
$ amtool --alertmanager.url=http://localhost:9093 silence query
ID Matchers Ends At Created By Comment
e534d465-1d58-21d0-270c-cd4c32cbcc4c alertname="TargetDown" instance="localhost:9100" 2026-09-22 22:02:33 UTC learner kernel upgrade, OPS-1234
$ amtool --alertmanager.url=http://localhost:9093 alert query
Alertname Starts At Summary State
$ amtool --alertmanager.url=http://localhost:9093 alert query -s
Alertname Starts At Summary State
TargetDown 2026-09-22 20:01:30 UTC node target localhost:9100 is down suppressed
These commands talk to the running Alertmanager, so they need --alertmanager.url= (where it listens). Piece by piece:
silence add+ matchers (alertname=TargetDown instance=localhost:9100) +-d 2h(duration) +-c '...'(comment) creates the silence and prints its ID, a long random identifier you use to refer to it later.silence querylists active silences: the ID, the matchers, when it ends, who created it and why.alert querylists the alerts Alertmanager holds. It hides silenced ones by default - hence the empty table.-s(--silenced) includes them; the TargetDown alert is there with Statesuppressed(received, but muted).
The rules for silences: always a comment with a ticket or reason, always the narrowest matchers that cover the work (not alertname=~".+", which matches every alert), always a duration, and expire it (amtool silence expire <id>) when you are done early. A forgotten broad silence is how real outages go unpaged.
Typing the URL every time gets old. Put it in ~/.config/amtool/config.yml once:
alertmanager.url: http://localhost:9093
Checking and reloading
$ amtool check-config /etc/prometheus/alertmanager.yml
Checking '/etc/prometheus/alertmanager.yml' SUCCESS
Found:
- global config
- route
- 1 inhibit rules
- 2 receivers
- 0 templates
amtool check-config validates a file without loading it: SUCCESS, then a count of what it found. (templates are optional files that customise the text of notifications; this box has none.)
A route pointing at a receiver name that is not in receivers: is the classic mistake. amtool prints FAILED: with the receiver's name and the route that uses it, then amtool: error: failed to validate 1 file(s) and exit code 1 - so it works as a gate in a CI pipeline.
sudo systemctl reload prometheus-alertmanager sends SIGHUP (2.21), which makes it re-read the file. Like Prometheus, a broken file is rejected and the old config keeps running; the journal says msg="Loading configuration file failed".
High availability, in one paragraph
If the one Alertmanager is down, nobody gets paged. So in production you run two or three and point every Prometheus at all of them. They talk to each other (a protocol called gossip, on port 9094) to share silences and to deduplicate notifications, so a page goes out once even though each Alertmanager received the alert. Do not put a load balancer (0.8, 9.23) in front of them for Prometheus to send to; the deduplication needs each one to see every alert.
Dashboard or alert?
If a human must act now, it is a page. If a human must act this week, it is a ticket. If it helps someone understand a problem they already know about, it belongs on a dashboard. CPU at 85%, a queue slightly longer than usual, GC pauses creeping up - dashboard. Checkout failing for users - page.
Later (Ch 29): you build those dashboards yourself, and add logs and traces for the "why" once a page has told you "what".
What you can now do
- Read an alertmanager.yml: the routing tree, grouping timers, receivers and inhibit rules.
- Predict and prove where an alert goes with
amtool config routes test. - Create, list and expire a narrow silence, and see silenced alerts with
alert query -s.