OnCallReady

Lesson 28.24 · Observability II: Alerting, Alertmanager & SLO alerts · 11 min read

The same rules in a cluster: PrometheusRule, cAdvisor, kube-state-metrics

In plain words

Imagine moving your family's fridge notes to a new house with a noticeboard system. The rules are the same: "if milk runs out, tell Mum". But now you do not stick them on the fridge; you pin them on the board in the hallway, and the house only reads notes with the family stamp on them. A note without the stamp is ignored, however important.

In Kubernetes with kube-prometheus-stack, rule files become PrometheusRule objects, scrape jobs become ServiceMonitors, and the operator loads only objects whose labels match its ruleSelector, often release: kube-prometheus-stack: that is the stamp. The metrics come from cAdvisor (what containers use), kube-state-metrics (what Kubernetes thinks) and the apps themselves.

The same rules, in a cluster

On a platform team your Prometheus runs inside Kubernetes, not on a VM, and somebody asks: "I added my alert, why does nothing happen?". Everything in this chapter carries over unchanged in substance. What changes is where the YAML lives and where the metrics come from.

What you need to know already: alerting rules and rule groups (28.1), promtool tests (28.4), Alertmanager config (28.7); the Prometheus Operator, kube-prometheus-stack and ServiceMonitor (27.4); Kubernetes objects, labels and selectors (15.26), Secrets (15.33), kubectl get ... -o jsonpath (15.38); custom resources such as Argo CD's Application (26.4); Helm charts (25.19); requests, limits, OOMKilled and throttling (17.1, 17.6); cgroup memory and exit 137 (5.11).

With the kube-prometheus-stack Helm chart (27.4: the Prometheus Operator plus Prometheus, Alertmanager, a dashboard tool, the node exporter and kube-state-metrics), you do not edit prometheus.yml or rule files on a disk:

apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
  name: orders-slo
  namespace: orders
  labels:
    release: kube-prometheus-stack     # the operator only picks up rules matching its ruleSelector
spec:
  groups:
    - name: orders-slo
      rules:
        - record: job:slo_errors_per_request:ratio_rate5m
          expr: |
            sum by (job) (rate(http_server_requests_seconds_count{job="orders", status=~"5.."}[5m]))
              /
            sum by (job) (rate(http_server_requests_seconds_count{job="orders"}[5m]))

Everything under spec: is the rule file from 28.14; the lines above it are the usual Kubernetes header (which API, which kind, name, namespace, labels).

The classic "my PrometheusRule does nothing" is the release: label. The operator is told which rule objects to load by a ruleSelector - a label selector (15.26) on the Prometheus custom resource. A rule object whose labels do not match it is silently ignored. Check what it wants with kubectl get prometheus -A -o jsonpath='{..ruleSelector}': -A looks in all namespaces, and the jsonpath {..ruleSelector} finds that field wherever it is nested.

promtool still works: extract spec to a plain rule file with yq (jq for YAML: yq '.spec' rule.yaml > rules.yml) and run promtool check rules and promtool test rules on it in CI.

Where cluster metrics come from

node exporter          machine metrics per node (what you used on oncall-lab)
kubelet / cAdvisor     per-container CPU, memory, filesystem, network:
                         container_cpu_usage_seconds_total
                         container_memory_working_set_bytes
kube-state-metrics     the STATE of Kubernetes objects, from the API server:
                         kube_pod_status_phase, kube_pod_container_status_restarts_total,
                         kube_deployment_status_replicas_available, kube_pod_container_resource_limits
the apps themselves    /metrics or /actuator/prometheus, via ServiceMonitors

Most useful alerts join the two.

Alerts every cluster has

The kube-prometheus-stack ships a few hundred rules. The ones worth knowing by heart, in simplified form:

# a container keeps crashing
- alert: KubePodCrashLooping
  expr: max_over_time(kube_pod_container_status_waiting_reason{reason="CrashLoopBackOff", job="kube-state-metrics"}[5m]) >= 1
  for: 15m

# a deployment is not at its desired replica count
- alert: KubeDeploymentReplicasMismatch
  expr: kube_deployment_spec_replicas != kube_deployment_status_replicas_available
  for: 15m

# a container was OOM-killed recently (chapter 5, in a cluster)
- alert: ContainerOOMKilled
  expr: increase(kube_pod_container_status_restarts_total[10m]) > 0
        and on (namespace, pod, container)
        kube_pod_container_status_last_terminated_reason{reason="OOMKilled"} == 1

# memory close to the limit: the next OOM is coming
- alert: ContainerMemoryNearLimit
  expr: |
    sum by (namespace, pod, container) (container_memory_working_set_bytes{container!=""})
      /
    sum by (namespace, pod, container) (kube_pod_container_resource_limits{resource="memory"})
      > 0.9
  for: 15m

Read them one at a time: the first keeps containers that were waiting in CrashLoopBackOff at any point in the last 5 minutes, for 15 minutes; the second keeps Deployments whose wanted replica count differs from the available one; the third keeps containers that restarted in the last 10 minutes and whose last exit was an OOM kill; the fourth divides memory used by the memory limit and keeps anything above 90%.

Look at the last two: they are and on (...) and a division between metrics from different exporters, matched on the labels they share - the vector matching of 27.15, doing real work. They are also cause alerts: route them to tickets or the team channel, and keep paging on the SLO burn.

Working set, not RSS

container_memory_working_set_bytes is the container's memory usage minus the inactive file cache (chapter 5: memory.current minus inactive_file) - memory the kernel cannot simply drop. It is the best predictor of the one thing that enforces the limit: the kernel's cgroup OOM killer at memory.max (exit 137, OOMKilled, 5.11). The kubelet does not evict a pod for crossing its own limit; eviction is node-level - the kubelet compares the node's memory.available (capacity minus the node's working set) with its eviction thresholds and then picks pods by usage above requests (17.3). container_memory_usage_bytes includes page cache that can be reclaimed and makes every JVM look like it is about to die. So: alert on working set against the limit (OOMKill risk), and on node memory pressure separately (eviction risk).

CPU throttling, the other quiet killer

sum by (namespace, pod, container) (rate(container_cpu_cfs_throttled_periods_total[5m]))
  /
sum by (namespace, pod, container) (rate(container_cpu_cfs_periods_total[5m]))

The kernel hands out a container's CPU limit in short periods (100 ms by default; CFS is the Linux CPU scheduler, 17.6). The first counter counts periods in which the container hit its limit and was paused; the second counts all periods. The ratio is the fraction of time slices in which it was throttled. A JVM with a 1-CPU limit, 20% average CPU and 40% throttled periods has p99 latency spikes that no average CPU graph explains - the interview question "why is p99 spiking at 40% average CPU" answered with a query. It is a dashboard and ticket signal; the page is the latency SLO it eventually burns.

What you can now do

Why it helps

Most of your alerting work in a real platform team will be in a cluster, not on a VM. "My PrometheusRule does nothing" is a weekly ticket, and it is almost always the missing release: label that the ruleSelector wants, silently ignored. You will also be asked to write the alerts every cluster needs: CrashLoopBackOff, replicas mismatch, OOMKilled, memory near limit, which means joining kube-state-metrics with cAdvisor using and on (namespace, pod, container).

Two signals come up in incidents and interviews. Working set, not memory usage, is what to compare with the limit, otherwise every JVM looks like it is about to die. And CPU throttling explains "p99 spikes while average CPU is 40%": a container hitting its CPU limit in short bursts. Knowing the query turns a mystery into a one-line answer.

FAQ

What is the difference between cAdvisor and kube-state-metrics?

cAdvisor, built into the kubelet, reports what containers actually use: CPU seconds, memory working set, filesystem and network, as container_* metrics. kube-state-metrics reads the Kubernetes API and reports the state of objects: pod phases, restarts, desired versus available replicas, requests and limits, as kube_* metrics. Useful alerts often combine them, for example memory used by cAdvisor divided by the limit from kube-state-metrics.

Why is my PrometheusRule not loaded?

Most often because its labels do not match the Prometheus resource's ruleSelector; with kube-prometheus-stack that usually means a label like release: kube-prometheus-stack. Also check ruleNamespaceSelector, which limits which namespaces are watched, and the operator logs for invalid rules. Inspect with kubectl get prometheus -A -o jsonpath='{..ruleSelector}' and check the rules page of the Prometheus UI.

Why use container_memory_working_set_bytes instead of container_memory_usage_bytes?

Usage includes page cache that the kernel can reclaim, so it creeps towards the limit on any container that reads files, and alerts on it are noise. Working set is usage minus inactive file cache, closer to what the cgroup OOM killer acts on and what the kubelet uses for memory pressure. Compare working set with the memory limit to predict OOM kills.

What does CPU throttling mean for a container?

With a CPU limit, the kernel's CFS quota allows the container a fixed amount of CPU time per period, usually 100 ms. If it uses its quota early in a period, it is paused until the next one. A multithreaded JVM can hit a 1-CPU limit in short bursts while average usage looks low, causing latency spikes. The ratio of container_cpu_cfs_throttled_periods_total to container_cpu_cfs_periods_total shows it.

Should I page on KubePodCrashLooping or ContainerOOMKilled?

Usually not directly. They are cause alerts: useful context, but a crash-looping pod in a Deployment with healthy replicas may not affect users, and a batch job can restart harmlessly. Route them to the team channel or a ticket, and page on the SLO burn of the user journey they eventually affect. Exceptions are single-replica critical components where one crash is an outage.

In an interview Mid

A service has p99 latency spikes while its average CPU is only 40% of its limit. What do you check?

CPU throttling. The kernel enforces a CPU limit in short periods (100 ms): a container that uses its quota early in a period is paused until the next one. Averages hide it - a JVM can average 20-40% of its limit and still be throttled in a large share of periods, which is exactly a p99 spike.

sum by (namespace, pod, container) (rate(container_cpu_cfs_throttled_periods_total[5m]))
  /
sum by (namespace, pod, container) (rate(container_cpu_cfs_periods_total[5m]))

= the fraction of periods in which the container was throttled. These container_* metrics come from cAdvisor in the kubelet (what containers use); kube_* from kube-state-metrics (what Kubernetes thinks: replicas, restarts, limits).

Fixes: raise or remove the CPU limit, or reduce bursts. Treat throttling as a dashboard/ticket signal; the page is the latency SLO it burns.

The memory equivalent: watch container_memory_working_set_bytes against the limit - that predicts the OOM kill (exit 137); container_memory_usage_bytes includes reclaimable cache.

Also asked: How do you monitor Kubernetes workloads with Prometheus? · What is the difference between cAdvisor and kube-state-metrics? · Your PrometheusRule object is not picked up by Prometheus. Why?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.