The restart storm
The problem. The most common probe outage is self-inflicted: a health check that is too strict or checks the wrong thing restarts every pod at once. This lesson shows how that happens and the few rules that prevent it - plus how to give slow-starting apps (Java ones especially) time to boot.
What you need to know already: the three probes and their fields (17.20), CrashLoopBackOff (15.14), CPU throttling (17.6), the JVM basics from 5.13.
A restart storm = many pods restarted over and over by their own health checks. Here is a config that looks responsible:
livenessProbe:
httpGet: {path: /actuator/health, port: 8080} # includes the database check
periodSeconds: 5
timeoutSeconds: 1
failureThreshold: 3
Monday, 09:00: the database gets slow - a big report query, 1.5s per health check query. Every ledger pod's /actuator/health now answers in 1.5s. The timeout is 1s, so the probe fails. Three failures, 15 seconds: the kubelet kills every ledger pod. They restart (Spring Boot: 45 seconds of CPU-heavy startup, no traffic served), immediately get probed against the still-slow database, get killed again. Restarts climb, CrashLoopBackOff back-off grows, the database gets hammered by connection pools reconnecting (a connection pool = a set of open database connections an app keeps ready and reuses). The database recovers at 09:10; the ledger does not, for another 15 minutes of back-off.
# an illustration: pay/ledger with the probes above (the probes mission)
kubectl get pods -n pay
NAME READY STATUS RESTARTS AGE
ledger-7b9c-kq2x8 0/1 CrashLoopBackOff 6 (82s ago) 3d
ledger-7b9c-pp3m1 0/1 Running 6 (12s ago) 3d
ledger-7b9c-z9x2c 0/1 CrashLoopBackOff 5 (2m ago) 3d
kubectl describe pod ledger-7b9c-kq2x8 -n pay | tail -4
Warning Unhealthy 3m (x18 over 12m) kubelet Liveness probe failed: Get "http://10.244.1.23:8080/actuator/health": context deadline exceeded (Client.Timeout exceeded while awaiting headers)
Normal Killing 3m (x6 over 12m) kubelet Container ledger failed liveness probe, will be restarted
Warning BackOff 82s (x19 over 9m) kubelet Back-off restarting failed container ledger in pod ledger-7b9c-kq2x8_pay(...)
A healthy service, taken down by its own health check. The rules that prevent it:
- Liveness must be cheap and local. It answers "is this process alive and able to make progress" - event loop not deadlocked, not out of heap, not stuck. It must never call the database, another service, or anything the pod cannot fix by restarting. Restarting never fixes somebody else.
- Readiness should reflect dependencies the pod truly needs to serve requests. Failing readiness removes the pod from traffic and costs nothing - no restart.
- Timeouts and thresholds must tolerate a bad minute.
timeoutSeconds: 1on a JVM that pauses for garbage collection (GC: the JVM freezing briefly to free unused memory) is asking for false positives; a liveness probe that kills after 15 seconds is aggressive for almost anything. - If in doubt, have no liveness probe. A crashing process restarts anyway (the container exits). Liveness is only for the process that is alive but stuck.
Spring Boot gives you exactly the right endpoints (Actuator = Spring Boot's built-in set of operational URLs, with management.endpoint.health.probes.enabled=true, the default on Kubernetes): /actuator/health/liveness (the application's LivenessState - local) and /actuator/health/readiness (ReadinessState, plus anything you add to the group, e.g. management.endpoint.health.group.readiness.include=readinessState,db). /actuator/health is the aggregate of everything and belongs on a dashboard, not in a probe.
Why startup probes exist
A JVM service needs time before it can answer anything: loading its code (class loading), building the app's objects (the Spring "context"), opening connection pools, and compiling hot code to machine code as it runs (JIT warm-up - Just-In-Time compilation). Under a CPU limit it is worse - startup is CPU-bound and throttled (part A): 45 seconds on a full core becomes 90 on limits.cpu: 500m.
Before startup probes existed (GA 1.20) you had two bad options:
- a big
initialDelaySecondson liveness (say 120s): every restart, even of a broken pod that crashed at second 3, waits two minutes before liveness helps; and if startup ever takes 121s, the kill loop starts; - a short delay: the kubelet kills the app while it is still starting - it never finishes, CrashLoopBackOff forever, exit code 143 each time.
Warning Unhealthy 1m (x9 over 2m) kubelet Liveness probe failed: Get "http://10.244.1.7:8080/actuator/health/liveness": dial tcp 10.244.1.7:8080: connect: connection refused
Normal Killing 1m (x3 over 2m) kubelet Container ledger failed liveness probe, will be restarted
connection refused every time = Tomcat (the web server built into a Spring Boot app) was not even listening on its port yet. That is the signature of a liveness probe firing during startup.
A startup probe separates the two phases: it may fail for a long time (the startup budget), and only once it succeeds do liveness (tight) and readiness start:
The Notion question: a Spring Boot app that takes 90 seconds to start
startupProbe: # budget: 30 x 5s = 150s, comfortably over 90s
httpGet: {path: /actuator/health/liveness, port: 8080}
periodSeconds: 5
failureThreshold: 30
readinessProbe: # traffic only when ready (and its deps are)
httpGet: {path: /actuator/health/readiness, port: 8080}
periodSeconds: 5
failureThreshold: 3
timeoutSeconds: 2
livenessProbe: # after startup: kill only if wedged for ~30s
httpGet: {path: /actuator/health/liveness, port: 8080}
periodSeconds: 10
failureThreshold: 3
timeoutSeconds: 2
How to justify every number in an interview:
- startup:
failureThreshold x periodSeconds> worst observed startup (90s) with margin, and a short period so the pod goes Ready soon after it is actually up. - readiness: fast reaction (5s) - it is cheap to be wrong, a pod just drops out of rotation briefly.
- liveness: slow reaction (30s+) - it is expensive to be wrong, a false positive restarts a healthy JVM. Local endpoint only.
timeoutSeconds: 2- a GC pause of more than a second must not count as dead.
And fix the cause, too: startup this slow under a 500m limit is the throttling from part A; giving the JVM more CPU (or no CPU limit) during startup often halves it.
Checklist for any probe you review
- Does liveness touch anything outside the process? -> move that check to readiness.
failureThreshold x periodSecondsfor liveness under ~20s? -> justify it or relax it.timeoutSecondsleft at 1 for a JVM or Python app? -> raise it.- Slow starter with no startup probe? -> add one sized to the startup time.
- Readiness identical to liveness? -> then one of them is pointless or wrong.
What you can now do
- Explain the restart storm and the rule "liveness is local, readiness may check dependencies".
- Size a startup probe:
failureThreshold x periodSeconds> worst startup time. - Review any probe config with the checklist above.