Chapter 28 Observability II: Alerting, Alertmanager & SLO alerts
Alerting rules and the for: clause, dead-target alerts, promtool rule tests, Alertmanager routing, grouping, inhibition and silences, and multi-window burn-rate alerts that actually fire.
In plain words
Imagine a smoke detector, a house phone and a family rule book. The detector decides that something is wrong: smoke above a level for long enough. The phone system decides who gets called: parents for a fire, the kids for "dinner is burnt", and if ten detectors go off at once, one call, not ten. The rule book says what counts as serious: a burnt toast once a week is fine, the kitchen on fire is not.
In Prometheus the detector is an alerting rule: a query plus for:. Alertmanager is the phone system: routing, grouping, inhibition and silences. The rule book is the SLO: 99.9% of orders requests succeed over 30 days, and burn-rate alerts page only when the error budget is being spent too fast. promtool test rules proves the detector actually works before the fire.
Why it matters on call
Being on call is mostly about alerts: which ones wake you, which ones you learn to ignore, and which ones stay silent during a real outage. This chapter is how you make sure a page means "users are hurting, act now". You will review teammates' rules for a for: 1h that hides 45-minute outages, a status=~"5xx" that can never match, an inhibit rule without equal that mutes everything, and routes pointing at receivers that do not exist.
SLO burn-rate alerting is a senior SRE interview staple: "why 14.4?", "why two windows?", "how fast does it detect a 5% outage?". And in Kubernetes the same rules live in PrometheusRule objects, where a missing release: label silently disables them, the kind of ticket a platform team gets every week. By the end you can write, test, route and justify alerts.
Lessons
- Alerting rules: for, labels, symptoms, dead targets
- promtool test rules: prove the alert fires
- Alertmanager: routing, grouping, inhibition, silences
- SLO alerts in Prometheus: multi-window, multi-burn-rate
- The same rules in a cluster: PrometheusRule, cAdvisor, kube-state-metrics
19 hands-on labs (missions, incidents and drills) run in the terminal: Open this chapter in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.
Questions people ask
What is the difference between an alert and a notification?
An alert is a Prometheus rule result: every series the expression returns is an alert instance, pending until for elapses and then firing. A notification is what Alertmanager sends: it groups firing alerts, routes them to a receiver like a pager or chat, throttles repeats, and applies silences and inhibitions. One notification can carry many alerts, and a firing alert may produce no notification at all if it is silenced or inhibited.
Should I alert on high CPU?
Usually not as a page. High CPU during a batch job hurts nobody, and a failing database may not move CPU at all. Page on symptoms users feel: error ratio, latency, availability, ideally as SLO burn rates. Keep cause-based signals such as CPU, memory near limit or throttling as tickets, chat messages or dashboard panels at low severity, because they help diagnose a problem you already know about.
What is an error budget?
The amount of unreliability the SLO allows: 1 minus the SLO. With 99.9% over 30 days, 0.1% of requests may fail, roughly 43 minutes of full outage per month. Teams spend it on releases and experiments; when it runs out, they prioritise reliability work. Burn rate measures how fast you consume it: 1 means you would use exactly the budget over the period, 14.4 means 2% of the month's budget in one hour.
How can I know an alert works if it has never fired?
Test it. promtool test rules evaluates your real rule files against synthetic series on a simulated clock and asserts which alerts fire when, including labels and annotations. Write both a positive case (it fires during a synthetic outage) and a negative case (it does not fire just before). Then, if possible, fire it for real in a test environment, for example with a load test that produces errors, and check the notification arrives.
Is Alertmanager required if my company already uses a paging service?
Yes, they do different jobs. Prometheus only decides that an alert fires and sends it to Alertmanager. Alertmanager does the routing, grouping, deduplication, inhibition and silences. A paging service such as PagerDuty or Opsgenie (a company that phones the on-call person, keeps the rota and escalates when nobody answers) sits after it, as one of Alertmanager's receivers. A common chain is Prometheus rules, then Alertmanager routing, then the paging service's escalation. The ideas are the same whichever tool does the routing.