OnCallReady

Lesson 17.20 · Kubernetes: Scheduling, Health & Security · 15 min read

Liveness, readiness, startup: what each answers and what failure does

In plain words

Imagine a shop with a manager who checks on each cashier all day. Three different questions. "Are you actually awake, or frozen staring at the wall?" If frozen, the manager sends them home and calls in a fresh one (liveness: restart). "Is your till open right now?" If not, customers are sent to other tills, but nobody is sent home (readiness: out of traffic). "Have you finished setting up your till this morning?" Until yes, the manager doesn't bother asking the other two questions (startup).

The kubelet on the node asks these questions with httpGet, tcpSocket, exec or grpc probes. A failed liveness or startup probe kills and restarts the container. A failed readiness probe only sets Ready=False, which removes the pod from Service endpoints.

Three questions the kubelet keeps asking

The problem. A process can be "running" and still useless: stuck in a loop, still starting up, or cut off from its database. Kubernetes cannot see that from outside unless you tell it how to check. The checks are called probes, and choosing the wrong one restarts healthy pods - or sends traffic to broken ones.

What you need to know already: pods and restarts (15.14), Services and their endpoints (16.1), systemd's Restart= (2.10), Docker's HEALTHCHECK (10.44), HTTP status codes (9.21), exit codes 137/143 (3.6).

A probe is a small health check - an HTTP request, a TCP connect or a command - that the kubelet runs against a container every few seconds.

The kubelet on the node runs every probe itself, against the container, forever (until the container stops). Each probe answers one question, and failing it has one consequence - memorise the table, it is the interview:

probequestionon failureon success
liveness"is this process wedged (stuck) beyond repair?"kubelet kills and restarts the container (restartPolicy)nothing
readiness"should this pod receive traffic right now?"pod is removed from Service endpoints (Ready=False). No restart.added back
startup"has the app finished starting?"after failureThreshold x periodSeconds: killed and restarted, like livenessliveness and readiness start running

Two consequences that people get wrong:

The fields

The paths below (/actuator/health/liveness) are health URLs that Spring Boot apps expose - Spring Boot is the Java web framework most of this course's services use. Any URL your app answers with 200 when healthy works the same way.

Later (Ch 21): you configure those Spring Boot health endpoints (Actuator) yourself.

livenessProbe:
  httpGet:
    path: /actuator/health/liveness
    port: 8080                 # a number or a named containerPort
  initialDelaySeconds: 0       # wait this long after container start before the first probe
  periodSeconds: 10            # probe every N seconds
  timeoutSeconds: 1            # a probe slower than this FAILS
  successThreshold: 1          # consecutive successes to count as passing (must be 1 for liveness/startup)
  failureThreshold: 3          # consecutive failures before acting
  terminationGracePeriodSeconds: 30   # optional: grace for a liveness/startup kill

The defaults are what you get when you write only the handler: period 10s, timeout 1s, failureThreshold 3, successThreshold 1, no initial delay. So a default liveness probe kills a container after ~30 seconds of failures (3 x 10s) - and after 1 second a slow response already counts as a failure.

Time to act = roughly initialDelaySeconds + failureThreshold x periodSeconds.

The four handlers

httpGet:   {path: /healthz, port: 8080}        # 200-399 = success
tcpSocket: {port: 5432}                        # the port accepts a connection = success
exec:      {command: [sh, -c, 'pg_isready -U postgres']}   # exit code 0 = success
grpc:      {port: 9090, service: ready}        # gRPC health checking protocol, SERVING = success

(pg_isready = the PostgreSQL database's own "are you accepting connections?" tool. gRPC = a way for programs to call each other over HTTP/2 with a binary format; it has a standard health-check call.)

httpGet is the kubelet itself making the request from the node, to the pod IP. exec forks a process inside the container every period - cheap for cat, expensive for a Java-based command-line tool; and it needs the binary to exist (distroless images, 10.38, have no shell).

What it looks like when they fail

kubectl describe pod shows the configuration...

    Liveness:     http-get http://:8080/actuator/health/liveness delay=0s timeout=1s period=10s #success=1 #failure=3
    Readiness:    http-get http://:8080/actuator/health/readiness delay=0s timeout=1s period=5s #success=1 #failure=3
    Startup:      http-get http://:8080/actuator/health/liveness delay=0s timeout=1s period=5s #success=1 #failure=30

...and the events tell you which one fired:

  Warning  Unhealthy  12s (x3 over 32s)  kubelet  Liveness probe failed: Get "http://10.244.1.23:8080/actuator/health/liveness": dial tcp 10.244.1.23:8080: connect: connection refused
  Normal   Killing    12s                kubelet  Container ledger failed liveness probe, will be restarted
  Warning  Unhealthy  4s (x9 over 44s)  kubelet  Readiness probe failed: HTTP probe failed with statuscode: 503
  Warning  Unhealthy  2s (x4 over 8s)   kubelet  Liveness probe failed: Get "http://10.244.1.23:8080/actuator/health": context deadline exceeded (Client.Timeout exceeded while awaiting headers)

The error text tells you why: connection refused = nothing listening yet (or the wrong port); statuscode: 503 = the app answered "not OK"; context deadline exceeded = it answered slower than timeoutSeconds.

A liveness kill shows up in the container status as a restart with exit code 143 (SIGTERM, if the app handled it and exited) or 137 (killed after the grace period, or a shell as PID 1 that ignores SIGTERM) and reason Error - not OOMKilled. Repeated kills become CrashLoopBackOff.

Readiness and the Service

# an illustration: pay/ledger with the probes above (the probes mission)
kubectl get pods -n pay -l app=ledger
NAME                 READY   STATUS    RESTARTS   AGE
ledger-7b9c-kq2x8    0/1     Running   0          6m
ledger-7b9c-pp3m1    0/1     Running   0          6m
kubectl get endpointslices -n pay -l kubernetes.io/service-name=ledger
NAME           ADDRESSTYPE   PORTS     ENDPOINTS   AGE
ledger-8q2zx   IPv4          8080      <unset>     6m

kubectl get endpointslices -l kubernetes.io/service-name=ledger lists the EndpointSlices (the lists of pod addresses behind a Service, 16.1) of the ledger Service. Running with 0/1 READY, zero restarts: readiness is failing. The EndpointSlice still lists the pods but with ready: false (-o yaml shows it), and kube-proxy sends them nothing. Rollouts use the same signal: a Deployment only counts a new pod available once it is Ready (plus minReadySeconds), so a failing readiness probe stops a bad rollout from replacing good pods - one of its most valuable effects.

Restart policy and probes

Probes act on containers; what happens after a liveness kill depends on restartPolicy: Always (Deployments) restarts, with the CrashLoopBackOff delay growing 10s, 20s, 40s ... up to 5 minutes and resetting after 10 minutes of running fine. Init containers cannot have probes (sidecars - init containers with restartPolicy: Always - can; 15.35).

What you can now do

Why it helps

Probes decide both restarts and traffic, so they're involved in a large share of incidents. A readiness probe failing on every replica is an outage with zero restarts: Running 0/1, empty endpoints. A liveness probe with the default 1-second timeout kills a GC-pausing JVM. A missing readiness probe lets a broken release replace every good pod in a minute, because "available" then just means "the process started".

In incident work, the event text is your diagnosis: connection refused (nothing listening yet or wrong port), statuscode: 503 (the app said no), context deadline exceeded (slower than timeoutSeconds). In interviews, the three-probe table is asked almost verbatim. In PR reviews you'll be the one checking thresholds and endpoints.

FAQ

Does a failing readiness probe restart the container?

No, never. It only marks the pod not Ready, which takes it out of every Service's endpoints so it gets no traffic. A pod can be Running 0/1 for days. If readiness fails on all replicas, the Service has no ready endpoints: an outage without a single restart.

What happens to liveness and readiness while a startup probe is running?

They aren't run at all until the startup probe succeeds once. That's the point of a startup probe: it gives a slow application a long budget (failureThreshold times periodSeconds) to start, after which the tight liveness and readiness checks take over. If the startup budget runs out, the container is killed like a liveness failure.

What are the probe defaults?

periodSeconds: 10, timeoutSeconds: 1, failureThreshold: 3, successThreshold: 1, no initial delay. So a default liveness probe restarts a container after about 30 seconds of failures, and a response slower than 1 second already counts as a failure. The 1-second timeout is the default most often worth changing.

What exit code does a liveness kill show?

Usually 143 (SIGTERM, when the app handles it and exits) or 137 (SIGKILL, when it's still running after the grace period, or a shell as PID 1 ignored SIGTERM), with reason Error, not OOMKilled. The events say "Container X failed liveness probe, will be restarted". Repeated kills turn into CrashLoopBackOff.

Where does an httpGet probe come from?

From the kubelet on the pod's node, which makes the HTTP request to the pod IP. It's not sent through the Service. That's why NetworkPolicies don't block probes (node traffic to its pods is always allowed) and why a probe can succeed while the Service path is broken, or the other way round.

In an interview Junior

Explain the difference between liveness, readiness and startup probes.

The kubelet runs all three against the container; each answers one question, and failing it has one consequence:

Handlers: httpGet (200-399), tcpSocket, exec, grpc. Defaults: period 10 s, timeout 1 s, failureThreshold 3 - time to act is about initialDelaySeconds + failureThreshold x periodSeconds.

The Unhealthy events say why: connection refused (not listening yet), statuscode: 503 (the app said no), context deadline exceeded (slower than the timeout). Readiness also gates rollouts: a new pod that never gets Ready stops a bad release.

Also asked: Pods are Running but 0/1 Ready and the Service returns errors. What is happening? · Which probe handler would you use for a database, a gRPC service and an HTTP API? · What exit code do you see after a liveness probe kills a container?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.