OnCallReady

Observability II: Alerting, Alertmanager & SLO alerts: interview questions

The question you are most likely to get for each topic, a model answer, and what else comes up. From chapter 28 of the course.

What makes a good alert? Mid

Every page should be urgent, actionable and about users. If the right response is "ignore it", delete the alert.

Also asked: Explain the multi-window, multi-burn-rate SLO alerting approach. · How would you reduce alert fatigue for an on-call team that gets too many pages? · How do you detect that a service stopped being monitored at all?

Explain the lifecycle of a Prometheus alert. Mid

An alerting rule is a PromQL expression evaluated every evaluation_interval; every series it returns is one alert instance (that is why up == 0 - a filter - returns exactly the dead targets). An empty result means all is fine.

Per instance:

You can watch it in /api/v1/alerts, and afterwards in the ALERTS series (alertstate pending or firing): max_over_time(ALERTS{alertname="X"}[3h]) tells you whether it ever fired.

Details that matter: labels must be static (a label with {{ $value }} creates a new alert each evaluation and for never completes); annotations are templates; "inactive, ok" in /api/v1/rules is also what a rule that can never match looks like.

Also asked: How do you detect that a service stopped being monitored at all? · What is the difference between symptom-based and cause-based alerting? · How do you choose the for duration of an alert?

Learn it: 28.1 Alerting rules: for, labels, symptoms, dead targets

How do you test Prometheus alerting rules? Mid

With promtool test rules: it evaluates the real rule files against synthetic series you write, on a simulated clock, with no running Prometheus.

A test file lists rule_files, then tests with input_series (a label set plus values in the shorthand - '1 1 1 0 0', '0+60x60' for a counter growing 60 per interval, _ for a missing sample) and alert_rule_test entries: at eval_time, exactly these firing alerts with these exp_labels and exp_annotations. promql_expr_test checks a recording rule's value.

What makes the tests worth having:

Run promtool check rules and test rules in the CI pipeline of the rules repo, so a broken alert cannot be merged.

Also asked: An alert rule passed code review but never fired during a real outage. How do you prevent this class of bug? · Why should every alert test include a negative case? · How would you check alert rules automatically before they are merged?

Learn it: 28.4 promtool test rules: prove the alert fires

During an outage the alert fired in Prometheus but the on-call engineer was not paged. How do you investigate? Mid

Follow the alert along its path, one hop at a time:

  1. Did it really fire? /api/v1/alerts or ALERTS{alertstate="firing"} - pending alerts never leave Prometheus.
  2. Did Alertmanager get it? amtool alert query - and -s, because silenced or inhibited alerts are hidden by default. State suppressed = received but muted.
  3. Was it muted? amtool silence query - a forgotten broad silence (alertname=~".+") is a classic. An inhibit rule without the right equal labels can mute everything.
  4. Where did it route? amtool config routes test severity=page job=orders prints the receiver. Routes are first-match: an earlier route may have swallowed it, or the severity label never said page.
  5. Timing: group_wait, group_interval and repeat_interval delay or batch notifications.
  6. Delivery: the receiver's own logs (the pager or webhook), and whether the last config reload failed - the journal says Loading configuration file failed while the old config keeps running.

Prevention: amtool check-config and routes test in CI, and two or three Alertmanagers that every Prometheus sends to.

Also asked: What does Alertmanager do? · What are grouping, inhibition and silences for in Alertmanager? · How would you design an Alertmanager routing tree for many teams?

Learn it: 28.7 Alertmanager: routing, grouping, inhibition, silences

Walk me through writing an SLO burn-rate alert for a 99.9% availability SLO. Mid

  1. SLI as a ratio of rates: 5xx over all requests, excluding health checks. Record it per window: job:slo_errors_per_request:ratio_rate5m, ..._rate1h, ..._rate30m, ..._rate6h.
  2. Budget = 1 - 0.999 = 0.1%. Burn rate = error ratio / 0.001; budget used = burn x window / 720 h.
  3. Conditions (the Workbook's table): page when the 1h and 5m ratios are both above 14.4 * 0.001 (2% of the month in an hour) or 6h and 30m above 6 * 0.001 (5%); ticket at 3 (1d/2h) and 1 (3d/6h).
(job:slo_errors_per_request:ratio_rate1h > (14.4 * 0.001)
   and job:slo_errors_per_request:ratio_rate5m > (14.4 * 0.001))
or
(job:slo_errors_per_request:ratio_rate6h > (6 * 0.001)
   and job:slo_errors_per_request:ratio_rate30m > (6 * 0.001))

Why two windows: the long one gives significance (a blip does not page), the short one makes it reset as soon as the problem stops. No for needed. Detection time scales with severity: a total outage pages in under a minute.

Watch: matching labels between windows (sum by (job) on both sides), low traffic (add a minimum-traffic clause), and test it with promtool.

Also asked: What are SLIs, SLOs and error budgets? · Why does each burn-rate condition use two windows? · How do SLO burn-rate alerts behave on a low-traffic service?

Learn it: 28.14 SLO alerts in Prometheus: multi-window, multi-burn-rate

A service has p99 latency spikes while its average CPU is only 40% of its limit. What do you check? Mid

CPU throttling. The kernel enforces a CPU limit in short periods (100 ms): a container that uses its quota early in a period is paused until the next one. Averages hide it - a JVM can average 20-40% of its limit and still be throttled in a large share of periods, which is exactly a p99 spike.

sum by (namespace, pod, container) (rate(container_cpu_cfs_throttled_periods_total[5m]))
  /
sum by (namespace, pod, container) (rate(container_cpu_cfs_periods_total[5m]))

= the fraction of periods in which the container was throttled. These container_* metrics come from cAdvisor in the kubelet (what containers use); kube_* from kube-state-metrics (what Kubernetes thinks: replicas, restarts, limits).

Fixes: raise or remove the CPU limit, or reduce bursts. Treat throttling as a dashboard/ticket signal; the page is the latency SLO it burns.

The memory equivalent: watch container_memory_working_set_bytes against the limit - that predicts the OOM kill (exit 137); container_memory_usage_bytes includes reclaimable cache.

Also asked: How do you monitor Kubernetes workloads with Prometheus? · What is the difference between cAdvisor and kube-state-metrics? · Your PrometheusRule object is not picked up by Prometheus. Why?

Learn it: 28.24 The same rules in a cluster: PrometheusRule, cAdvisor, kube-state-metrics

Practise these answers with flashcards and labs Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.