The same rules, in a cluster
On a platform team your Prometheus runs inside Kubernetes, not on a VM, and somebody asks: "I added my alert, why does nothing happen?". Everything in this chapter carries over unchanged in substance. What changes is where the YAML lives and where the metrics come from.
What you need to know already: alerting rules and rule groups (28.1), promtool tests (28.4), Alertmanager config (28.7); the Prometheus Operator, kube-prometheus-stack and ServiceMonitor (27.4); Kubernetes objects, labels and selectors (15.26), Secrets (15.33), kubectl get ... -o jsonpath (15.38); custom resources such as Argo CD's Application (26.4); Helm charts (25.19); requests, limits, OOMKilled and throttling (17.1, 17.6); cgroup memory and exit 137 (5.11).
With the kube-prometheus-stack Helm chart (27.4: the Prometheus Operator plus Prometheus, Alertmanager, a dashboard tool, the node exporter and kube-state-metrics), you do not edit prometheus.yml or rule files on a disk:
- scrape targets are ServiceMonitor / PodMonitor objects (27.4);
- rules are PrometheusRule objects - a custom resource whose
specis exactly the rule-file format you have been writing; - Alertmanager's config is a Secret (or AlertmanagerConfig objects, one per namespace, which the operator merges into the routing tree).
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: orders-slo
namespace: orders
labels:
release: kube-prometheus-stack # the operator only picks up rules matching its ruleSelector
spec:
groups:
- name: orders-slo
rules:
- record: job:slo_errors_per_request:ratio_rate5m
expr: |
sum by (job) (rate(http_server_requests_seconds_count{job="orders", status=~"5.."}[5m]))
/
sum by (job) (rate(http_server_requests_seconds_count{job="orders"}[5m]))
Everything under spec: is the rule file from 28.14; the lines above it are the usual Kubernetes header (which API, which kind, name, namespace, labels).
The classic "my PrometheusRule does nothing" is the release: label. The operator is told which rule objects to load by a ruleSelector - a label selector (15.26) on the Prometheus custom resource. A rule object whose labels do not match it is silently ignored. Check what it wants with kubectl get prometheus -A -o jsonpath='{..ruleSelector}': -A looks in all namespaces, and the jsonpath {..ruleSelector} finds that field wherever it is nested.
promtool still works: extract spec to a plain rule file with yq (jq for YAML: yq '.spec' rule.yaml > rules.yml) and run promtool check rules and promtool test rules on it in CI.
Where cluster metrics come from
node exporter machine metrics per node (what you used on oncall-lab)
kubelet / cAdvisor per-container CPU, memory, filesystem, network:
container_cpu_usage_seconds_total
container_memory_working_set_bytes
kube-state-metrics the STATE of Kubernetes objects, from the API server:
kube_pod_status_phase, kube_pod_container_status_restarts_total,
kube_deployment_status_replicas_available, kube_pod_container_resource_limits
the apps themselves /metrics or /actuator/prometheus, via ServiceMonitors
- cAdvisor ("container advisor") is built into the kubelet (15.7). It reads each container's cgroup and exports what the container uses, as
container_*metrics. - kube-state-metrics is a small exporter that watches the Kubernetes API and exports what Kubernetes thinks: desired vs actual replicas, restarts, phases, requests and limits, as
kube_*metrics.
Most useful alerts join the two.
Alerts every cluster has
The kube-prometheus-stack ships a few hundred rules. The ones worth knowing by heart, in simplified form:
# a container keeps crashing
- alert: KubePodCrashLooping
expr: max_over_time(kube_pod_container_status_waiting_reason{reason="CrashLoopBackOff", job="kube-state-metrics"}[5m]) >= 1
for: 15m
# a deployment is not at its desired replica count
- alert: KubeDeploymentReplicasMismatch
expr: kube_deployment_spec_replicas != kube_deployment_status_replicas_available
for: 15m
# a container was OOM-killed recently (chapter 5, in a cluster)
- alert: ContainerOOMKilled
expr: increase(kube_pod_container_status_restarts_total[10m]) > 0
and on (namespace, pod, container)
kube_pod_container_status_last_terminated_reason{reason="OOMKilled"} == 1
# memory close to the limit: the next OOM is coming
- alert: ContainerMemoryNearLimit
expr: |
sum by (namespace, pod, container) (container_memory_working_set_bytes{container!=""})
/
sum by (namespace, pod, container) (kube_pod_container_resource_limits{resource="memory"})
> 0.9
for: 15m
Read them one at a time: the first keeps containers that were waiting in CrashLoopBackOff at any point in the last 5 minutes, for 15 minutes; the second keeps Deployments whose wanted replica count differs from the available one; the third keeps containers that restarted in the last 10 minutes and whose last exit was an OOM kill; the fourth divides memory used by the memory limit and keeps anything above 90%.
Look at the last two: they are and on (...) and a division between metrics from different exporters, matched on the labels they share - the vector matching of 27.15, doing real work. They are also cause alerts: route them to tickets or the team channel, and keep paging on the SLO burn.
Working set, not RSS
container_memory_working_set_bytes is the container's memory usage minus the inactive file cache (chapter 5: memory.current minus inactive_file) - memory the kernel cannot simply drop. It is the best predictor of the one thing that enforces the limit: the kernel's cgroup OOM killer at memory.max (exit 137, OOMKilled, 5.11). The kubelet does not evict a pod for crossing its own limit; eviction is node-level - the kubelet compares the node's memory.available (capacity minus the node's working set) with its eviction thresholds and then picks pods by usage above requests (17.3). container_memory_usage_bytes includes page cache that can be reclaimed and makes every JVM look like it is about to die. So: alert on working set against the limit (OOMKill risk), and on node memory pressure separately (eviction risk).
CPU throttling, the other quiet killer
sum by (namespace, pod, container) (rate(container_cpu_cfs_throttled_periods_total[5m]))
/
sum by (namespace, pod, container) (rate(container_cpu_cfs_periods_total[5m]))
The kernel hands out a container's CPU limit in short periods (100 ms by default; CFS is the Linux CPU scheduler, 17.6). The first counter counts periods in which the container hit its limit and was paused; the second counts all periods. The ratio is the fraction of time slices in which it was throttled. A JVM with a 1-CPU limit, 20% average CPU and 40% throttled periods has p99 latency spikes that no average CPU graph explains - the interview question "why is p99 spiking at 40% average CPU" answered with a query. It is a dashboard and ticket signal; the page is the latency SLO it eventually burns.
What you can now do
- Ship a rule as a PrometheusRule and find out why the operator ignores it.
- Say which metrics come from cAdvisor and which from kube-state-metrics, and join them in an alert.
- Alert on memory near the limit and on CPU throttling without paging on them.