Observability II: Alerting, Alertmanager & SLO alerts: interview questions
The question you are most likely to get for each topic, a model answer, and what else comes up. From chapter 28 of the course.
What makes a good alert? Mid
Every page should be urgent, actionable and about users. If the right response is "ignore it", delete the alert.
- Page on symptoms, not causes: errors, latency, availability as users see them. CPU at 85% or a growing queue are tickets or dashboard material.
- Tie it to the SLO: a multi-window burn-rate alert (e.g. 1h and 5m windows above burn rate 14.4 = 2% of the monthly budget in an hour) pages in proportion to how bad things are and resets fast.
- A sensible
for- long enough to ignore a blip, never longer than the outages you care about; static labels (severity), templated annotations with a summary, dashboard link andrunbook_url. - Cover the silent failures:
up == 0for dead targets andabsent(...)for jobs or metrics that vanished. - Prove it fires:
promtool test ruleswith positive and negative cases, in CI. - Deliver it well: Alertmanager routes pages and tickets to different receivers, groups related alerts into one notification, inhibits redundant ones, and silences are narrow, commented and time-limited.
Also asked: Explain the multi-window, multi-burn-rate SLO alerting approach. · How would you reduce alert fatigue for an on-call team that gets too many pages? · How do you detect that a service stopped being monitored at all?
Explain the lifecycle of a Prometheus alert. Mid
An alerting rule is a PromQL expression evaluated every evaluation_interval; every series it returns is one alert instance (that is why up == 0 - a filter - returns exactly the dead targets). An empty result means all is fine.
Per instance:
- inactive - the expression does not return it.
- pending - it is returned, but not yet continuously for the rule's
for. The clock starts atactiveAt; one false evaluation resets it. - firing - returned for at least
for. Only now is it sent to Alertmanager, which notifies people. - When the expression stops returning it, the alert resolves (
keep_firing_forcan hold it a little longer against flapping).
You can watch it in /api/v1/alerts, and afterwards in the ALERTS series (alertstate pending or firing): max_over_time(ALERTS{alertname="X"}[3h]) tells you whether it ever fired.
Details that matter: labels must be static (a label with {{ $value }} creates a new alert each evaluation and for never completes); annotations are templates; "inactive, ok" in /api/v1/rules is also what a rule that can never match looks like.
Also asked: How do you detect that a service stopped being monitored at all? · What is the difference between symptom-based and cause-based alerting? · How do you choose the for duration of an alert?
Learn it: 28.1 Alerting rules: for, labels, symptoms, dead targets
How do you test Prometheus alerting rules? Mid
With promtool test rules: it evaluates the real rule files against synthetic series you write, on a simulated clock, with no running Prometheus.
A test file lists rule_files, then tests with input_series (a label set plus values in the shorthand - '1 1 1 0 0', '0+60x60' for a counter growing 60 per interval, _ for a missing sample) and alert_rule_test entries: at eval_time, exactly these firing alerts with these exp_labels and exp_annotations. promql_expr_test checks a recording rule's value.
What makes the tests worth having:
- a positive case (it fires when it should) and a negative case (
exp_alerts: []just before) - a test that only checks "it fires" passes for an alert that always fires; - the
forarithmetic: down from minute 3 withfor: 2mfires at 5m; - the edges: a gap (missing is not zero), a counter reset, no traffic (0/0 = NaN never fires);
- labels and annotations compared exactly - a wrong severity fails the test.
Run promtool check rules and test rules in the CI pipeline of the rules repo, so a broken alert cannot be merged.
Also asked: An alert rule passed code review but never fired during a real outage. How do you prevent this class of bug? · Why should every alert test include a negative case? · How would you check alert rules automatically before they are merged?
During an outage the alert fired in Prometheus but the on-call engineer was not paged. How do you investigate? Mid
Follow the alert along its path, one hop at a time:
- Did it really fire?
/api/v1/alertsorALERTS{alertstate="firing"}- pending alerts never leave Prometheus. - Did Alertmanager get it?
amtool alert query- and-s, because silenced or inhibited alerts are hidden by default. Statesuppressed= received but muted. - Was it muted?
amtool silence query- a forgotten broad silence (alertname=~".+") is a classic. An inhibit rule without the rightequallabels can mute everything. - Where did it route?
amtool config routes test severity=page job=ordersprints the receiver. Routes are first-match: an earlier route may have swallowed it, or theseveritylabel never saidpage. - Timing:
group_wait,group_intervalandrepeat_intervaldelay or batch notifications. - Delivery: the receiver's own logs (the pager or webhook), and whether the last config reload failed - the journal says
Loading configuration file failedwhile the old config keeps running.
Prevention: amtool check-config and routes test in CI, and two or three Alertmanagers that every Prometheus sends to.
Also asked: What does Alertmanager do? · What are grouping, inhibition and silences for in Alertmanager? · How would you design an Alertmanager routing tree for many teams?
Learn it: 28.7 Alertmanager: routing, grouping, inhibition, silences
Walk me through writing an SLO burn-rate alert for a 99.9% availability SLO. Mid
- SLI as a ratio of rates: 5xx over all requests, excluding health checks. Record it per window:
job:slo_errors_per_request:ratio_rate5m,..._rate1h,..._rate30m,..._rate6h. - Budget = 1 - 0.999 = 0.1%. Burn rate = error ratio / 0.001; budget used = burn x window / 720 h.
- Conditions (the Workbook's table): page when the 1h and 5m ratios are both above
14.4 * 0.001(2% of the month in an hour) or 6h and 30m above6 * 0.001(5%); ticket at 3 (1d/2h) and 1 (3d/6h).
(job:slo_errors_per_request:ratio_rate1h > (14.4 * 0.001)
and job:slo_errors_per_request:ratio_rate5m > (14.4 * 0.001))
or
(job:slo_errors_per_request:ratio_rate6h > (6 * 0.001)
and job:slo_errors_per_request:ratio_rate30m > (6 * 0.001))
Why two windows: the long one gives significance (a blip does not page), the short one makes it reset as soon as the problem stops. No for needed. Detection time scales with severity: a total outage pages in under a minute.
Watch: matching labels between windows (sum by (job) on both sides), low traffic (add a minimum-traffic clause), and test it with promtool.
Also asked: What are SLIs, SLOs and error budgets? · Why does each burn-rate condition use two windows? · How do SLO burn-rate alerts behave on a low-traffic service?
Learn it: 28.14 SLO alerts in Prometheus: multi-window, multi-burn-rate
A service has p99 latency spikes while its average CPU is only 40% of its limit. What do you check? Mid
CPU throttling. The kernel enforces a CPU limit in short periods (100 ms): a container that uses its quota early in a period is paused until the next one. Averages hide it - a JVM can average 20-40% of its limit and still be throttled in a large share of periods, which is exactly a p99 spike.
sum by (namespace, pod, container) (rate(container_cpu_cfs_throttled_periods_total[5m]))
/
sum by (namespace, pod, container) (rate(container_cpu_cfs_periods_total[5m]))
= the fraction of periods in which the container was throttled. These container_* metrics come from cAdvisor in the kubelet (what containers use); kube_* from kube-state-metrics (what Kubernetes thinks: replicas, restarts, limits).
Fixes: raise or remove the CPU limit, or reduce bursts. Treat throttling as a dashboard/ticket signal; the page is the latency SLO it burns.
The memory equivalent: watch container_memory_working_set_bytes against the limit - that predicts the OOM kill (exit 137); container_memory_usage_bytes includes reclaimable cache.
Also asked: How do you monitor Kubernetes workloads with Prometheus? · What is the difference between cAdvisor and kube-state-metrics? · Your PrometheusRule object is not picked up by Prometheus. Why?
Learn it: 28.24 The same rules in a cluster: PrometheusRule, cAdvisor, kube-state-metrics
Practise these answers with flashcards and labs Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.