OnCallReady

Lesson 17.22 · Kubernetes: Scheduling, Health & Security · 11 min read

Probe design: the restart storm, and why JVMs need startup probes

In plain words

Imagine a smoke alarm wired so that it also goes off whenever the kitchen next door is busy. On a busy Monday, it rings in every flat, everyone evacuates, comes back, it rings again, and the building is empty all morning even though nothing was on fire. The alarm was supposed to detect a problem in your flat, not in the kitchen next door.

That's the restart storm: a liveness probe that checks the database kills every healthy pod when the database gets slow, and the restarts make it worse. Liveness must be cheap and local; dependencies belong in readiness, which only takes a pod out of traffic. And a slow starter, like a JVM, needs a startup probe so liveness doesn't kill it before it has finished starting.

The restart storm

The problem. The most common probe outage is self-inflicted: a health check that is too strict or checks the wrong thing restarts every pod at once. This lesson shows how that happens and the few rules that prevent it - plus how to give slow-starting apps (Java ones especially) time to boot.

What you need to know already: the three probes and their fields (17.20), CrashLoopBackOff (15.14), CPU throttling (17.6), the JVM basics from 5.13.

A restart storm = many pods restarted over and over by their own health checks. Here is a config that looks responsible:

livenessProbe:
  httpGet: {path: /actuator/health, port: 8080}   # includes the database check
  periodSeconds: 5
  timeoutSeconds: 1
  failureThreshold: 3

Monday, 09:00: the database gets slow - a big report query, 1.5s per health check query. Every ledger pod's /actuator/health now answers in 1.5s. The timeout is 1s, so the probe fails. Three failures, 15 seconds: the kubelet kills every ledger pod. They restart (Spring Boot: 45 seconds of CPU-heavy startup, no traffic served), immediately get probed against the still-slow database, get killed again. Restarts climb, CrashLoopBackOff back-off grows, the database gets hammered by connection pools reconnecting (a connection pool = a set of open database connections an app keeps ready and reuses). The database recovers at 09:10; the ledger does not, for another 15 minutes of back-off.

# an illustration: pay/ledger with the probes above (the probes mission)
kubectl get pods -n pay
NAME                 READY   STATUS             RESTARTS       AGE
ledger-7b9c-kq2x8    0/1     CrashLoopBackOff   6 (82s ago)    3d
ledger-7b9c-pp3m1    0/1     Running            6 (12s ago)    3d
ledger-7b9c-z9x2c    0/1     CrashLoopBackOff   5 (2m ago)     3d
kubectl describe pod ledger-7b9c-kq2x8 -n pay | tail -4
  Warning  Unhealthy  3m (x18 over 12m)  kubelet  Liveness probe failed: Get "http://10.244.1.23:8080/actuator/health": context deadline exceeded (Client.Timeout exceeded while awaiting headers)
  Normal   Killing    3m (x6 over 12m)   kubelet  Container ledger failed liveness probe, will be restarted
  Warning  BackOff    82s (x19 over 9m)  kubelet  Back-off restarting failed container ledger in pod ledger-7b9c-kq2x8_pay(...)

A healthy service, taken down by its own health check. The rules that prevent it:

  1. Liveness must be cheap and local. It answers "is this process alive and able to make progress" - event loop not deadlocked, not out of heap, not stuck. It must never call the database, another service, or anything the pod cannot fix by restarting. Restarting never fixes somebody else.
  2. Readiness should reflect dependencies the pod truly needs to serve requests. Failing readiness removes the pod from traffic and costs nothing - no restart.
  3. Timeouts and thresholds must tolerate a bad minute. timeoutSeconds: 1 on a JVM that pauses for garbage collection (GC: the JVM freezing briefly to free unused memory) is asking for false positives; a liveness probe that kills after 15 seconds is aggressive for almost anything.
  4. If in doubt, have no liveness probe. A crashing process restarts anyway (the container exits). Liveness is only for the process that is alive but stuck.

Spring Boot gives you exactly the right endpoints (Actuator = Spring Boot's built-in set of operational URLs, with management.endpoint.health.probes.enabled=true, the default on Kubernetes): /actuator/health/liveness (the application's LivenessState - local) and /actuator/health/readiness (ReadinessState, plus anything you add to the group, e.g. management.endpoint.health.group.readiness.include=readinessState,db). /actuator/health is the aggregate of everything and belongs on a dashboard, not in a probe.

Why startup probes exist

A JVM service needs time before it can answer anything: loading its code (class loading), building the app's objects (the Spring "context"), opening connection pools, and compiling hot code to machine code as it runs (JIT warm-up - Just-In-Time compilation). Under a CPU limit it is worse - startup is CPU-bound and throttled (part A): 45 seconds on a full core becomes 90 on limits.cpu: 500m.

Before startup probes existed (GA 1.20) you had two bad options:

  Warning  Unhealthy  1m (x9 over 2m)  kubelet  Liveness probe failed: Get "http://10.244.1.7:8080/actuator/health/liveness": dial tcp 10.244.1.7:8080: connect: connection refused
  Normal   Killing    1m (x3 over 2m)  kubelet  Container ledger failed liveness probe, will be restarted

connection refused every time = Tomcat (the web server built into a Spring Boot app) was not even listening on its port yet. That is the signature of a liveness probe firing during startup.

A startup probe separates the two phases: it may fail for a long time (the startup budget), and only once it succeeds do liveness (tight) and readiness start:

The Notion question: a Spring Boot app that takes 90 seconds to start

startupProbe:                        # budget: 30 x 5s = 150s, comfortably over 90s
  httpGet: {path: /actuator/health/liveness, port: 8080}
  periodSeconds: 5
  failureThreshold: 30
readinessProbe:                      # traffic only when ready (and its deps are)
  httpGet: {path: /actuator/health/readiness, port: 8080}
  periodSeconds: 5
  failureThreshold: 3
  timeoutSeconds: 2
livenessProbe:                       # after startup: kill only if wedged for ~30s
  httpGet: {path: /actuator/health/liveness, port: 8080}
  periodSeconds: 10
  failureThreshold: 3
  timeoutSeconds: 2

How to justify every number in an interview:

And fix the cause, too: startup this slow under a 500m limit is the throttling from part A; giving the JVM more CPU (or no CPU limit) during startup often halves it.

Checklist for any probe you review

What you can now do

Why it helps

A health check taking down a healthy service is one of the most common self-inflicted outages in Kubernetes, and one of the most embarrassing postmortems. The config that causes it looks responsible, so it passes review unless someone knows the pattern. That someone should be you.

On Java platforms, startup probes are the difference between a 90-second Spring Boot app that deploys cleanly and one that crash-loops forever with connection refused and exit 143, especially under a CPU limit that doubles startup time. "Design probes for a Spring Boot app that takes 90 seconds to start" is a Notion question and a common interview task; this lesson gives you numbers you can justify line by line.

FAQ

Why shouldn't liveness check the database?

Because restarting your pod can't fix the database. If the database gets slow, every pod's liveness check fails at once, the kubelet kills them all, they restart, reconnect in a burst (hammering the database further), get probed against the still-slow database and are killed again. A dependency check belongs in readiness, which removes the pod from traffic without restarting it.

Is it OK to have no liveness probe at all?

Often, yes. A process that crashes exits, and the container restarts anyway. Liveness is only useful for a process that's alive but stuck: deadlocked, wedged, out of heap but not exiting. If you can't define a cheap, local check for that, no liveness probe is safer than a bad one.

How do I size a startup probe?

failureThreshold times periodSeconds must be comfortably above the worst observed startup time, including under CPU limits and on busy nodes. For a 90-second start: periodSeconds: 5 and failureThreshold: 30 gives 150 seconds. A short period means the pod goes Ready soon after it's actually up; the long threshold is the budget.

Why does my app crash-loop with "connection refused" on liveness right after deploy?

The liveness probe is firing during startup, before the server is listening. Without a startup probe, liveness starts immediately (or after initialDelaySeconds), fails three times and kills the app mid-start, forever. Add a startup probe with a budget longer than startup; liveness only begins once it succeeds.

Why does startup take twice as long in Kubernetes as on my laptop?

Often CPU throttling. JVM startup is CPU-heavy (class loading, JIT, Spring context), and a limits.cpu: 500m gives it half a core per 100 ms period. A 45-second start on a full core becomes about 90 seconds. Raising or removing the CPU limit, or giving it more CPU, often halves startup, and your startup probe budget must account for it.

In an interview Mid

How would you design the probes for a Java service that takes 90 seconds to start?

Three probes, three jobs, each number justified:

The reason: a liveness probe that checks the database turns a slow database into a restart storm - every pod killed at once, each restart a 90-second boot. Also check the CPU limit: startup is CPU-bound, and throttling can double it.

Also asked: What is a restart storm, and how do you prevent one? · Why must a liveness probe never check a dependency? · How would you review the probes in a teammate's Deployment YAML?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.