OnCallReady

Lesson 21.3 · Spring Boot Runtime, Resilience & Python Ops · 16 min read

Health groups, liveness vs readiness, and the three probes

In plain words

Imagine a lifeguard checking two different things. First: "Is this swimmer conscious?" If not, pull them out right away. Second: "Is the pool ready for more swimmers?" If the water is being cleaned, you just stop letting people in for a while, and nobody gets pulled out.

Kubernetes asks a service the same two questions. The liveness probe is "are you conscious?", and a failure makes the kubelet restart the container. The readiness probe is "can you take visitors?", and a failure just stops traffic to that pod. Spring Boot answers them at /actuator/health/liveness and /actuator/health/readiness. The key rule: a database outage is a "stop letting people in" problem, never a reason to pull every swimmer out.

/actuator/health

The problem. Lesson 17.22 said "liveness must be local, readiness may check dependencies". In a Spring Boot app that choice is a few properties - and the defaults change depending on whether the app knows it runs in Kubernetes.

What you need to know already: Actuator (21.1), the three probes and the restart storm (17.20, 17.22), HTTP status codes (9.21).

$ curl -s localhost:8080/actuator/health
{"status":"UP"}

By default you only see the aggregate. With details:

management.endpoint.health.show-details=always          # or when-authorized
$ curl -s localhost:8080/actuator/health | jq .
{
  "status": "UP",
  "components": {
    "db":        { "status": "UP", "details": { "database": "PostgreSQL", "validationQuery": "isValid()" } },
    "diskSpace": { "status": "UP", "details": { "total": 19338280960, "free": 11524853760, "threshold": 10485760, ... } },
    "livenessState":  { "status": "UP" },
    "ping":      { "status": "UP" },
    "readinessState": { "status": "UP" },
    "ssl":       { "status": "UP", ... }
  },
  "groups": [ "liveness", "readiness" ]
}

Every HealthIndicator (a small built-in check) on the classpath (the set of libraries the app was built with - like its node_modules) contributes a component: db (a database connection, a "DataSource", is configured), diskSpace, redis, rabbit, kafka (caches and message queues)... The aggregate is DOWN if any component is DOWN - which is exactly why the aggregate must not be your liveness probe.

The HTTP status code carries the result: 200 for UP, 503 for DOWN or OUT_OF_SERVICE. Probes look at the status code, not the body.

The two groups

management.endpoint.health.probes.enabled=true

gives you two health groups (a group = a named subset of the components, with its own URL):

/actuator/health/liveness     livenessState:  CORRECT | BROKEN       -> UP | DOWN
/actuator/health/readiness    readinessState: ACCEPTING_TRAFFIC | REFUSING_TRAFFIC -> UP | OUT_OF_SERVICE

Boot enables them automatically when it detects it is running on Kubernetes (it sees the KUBERNETES_SERVICE_HOST variables). On a VM, in a test, or in Docker locally they are off until you set the property - and the paths 404:

# before probes.enabled is set (the actuator mission set it: yours answers UP now)
curl -s localhost:8080/actuator/health/liveness
{"timestamp":"...","status":404,"error":"Not Found","path":"/actuator/health/liveness"}

The states are driven by the application lifecycle: readiness becomes ACCEPTING_TRAFFIC only once the app has started (after its startup tasks, "runners"), and flips to REFUSING_TRAFFIC the moment graceful shutdown begins. Liveness goes BROKEN if the application publishes it (a fatal internal error).

Liveness vs readiness: the rule

liveness   "restart me"            must be cheap, LOCAL, and only fail if a restart would fix it
readiness  "stop sending me traffic"  may reflect dependencies the app cannot work without
management.endpoint.health.group.readiness.include=readinessState,db
management.endpoint.health.group.liveness.include=livenessState        # the default; keep it

But think about it: if every pod goes unready at once, the Service has no endpoints and callers get connection errors instead of a fast 503 from you. Many teams keep readiness local too and handle dependency failure in code (circuit breaker, fallback).

Running management on a separate port

management.server.port=8081

Separate port, separate (small) thread pool (a fixed set of threads that take turns doing the work, 20.20): a probe still answers when all 200 Tomcat request threads are stuck. With a separate port the probe paths also exist on the main port if you set management.endpoint.health.probes.add-additional-paths=true (/livez and /readyz).

The three probes for an app that takes 90 seconds to start

startupProbe:
  httpGet: { path: /actuator/health/liveness, port: management }
  periodSeconds: 5
  failureThreshold: 30          # 30 x 5 s = 150 s allowed to start: 90 s plus margin
livenessProbe:
  httpGet: { path: /actuator/health/liveness, port: management }
  periodSeconds: 10
  timeoutSeconds: 2
  failureThreshold: 3           # 30 s of consecutive failure before a restart
readinessProbe:
  httpGet: { path: /actuator/health/readiness, port: management }
  periodSeconds: 5
  timeoutSeconds: 2
  failureThreshold: 2

Reading the status code from the shell

$ curl -s -o /dev/null -w '%{http_code}\n' localhost:8080/actuator/health/readiness
200
$ curl -sf localhost:8080/actuator/health/liveness >/dev/null && echo alive || echo dead
alive

-f makes curl exit non-zero on 4xx/5xx - that is what a shell probe (or a systemd watchdog script) relies on. -o /dev/null -w '%{http_code}\n' = throw the body away and print only the status code.

What you can now do

Why it helps

Badly designed probes cause some of the worst self-inflicted outages. Liveness pointed at the aggregate /actuator/health, which includes db, means a 30-second database blip restarts every pod at once, and they all come back hammering the recovering database. A 1-second probe timeout on a service that sometimes pauses for GC means restarts under load, exactly when you need capacity.

These are the settings you will review in every Deployment: which path, which port, timeoutSeconds, failureThreshold, and whether a slow-starting service has a startup probe. It is also a CKA topic and one of the most common platform interview questions, and you'll be able to answer it with the specific Spring Boot endpoints behind it.

Commands in this lesson

curl

FAQ

What does the startup probe add?

It holds off liveness and readiness until it succeeds once. For an app that takes 90 seconds to start, a startup probe with periodSeconds: 5 and failureThreshold: 30 gives it up to 150 seconds. Without it you'd need initialDelaySeconds: 100 on liveness, which delays detecting real hangs forever after, or liveness kills the pod mid-startup and a healthy app ends in CrashLoopBackOff.

Should readiness include the database?

It may, but think about it. If readiness includes db and the database goes down, every pod becomes unready at once, the Service has no endpoints, and callers get connection errors instead of a fast 503 from your app. Many teams keep readiness local too and handle dependency failure in code with circuit breakers and fallbacks. What you must never do is put the database in liveness.

Why are /actuator/health/liveness and /readiness 404 on my VM?

Spring Boot only enables the probe health groups automatically when it detects Kubernetes, through the KUBERNETES_SERVICE_HOST environment variables. On a VM, in Docker locally or in tests they are off, and the paths return 404. Set management.endpoint.health.probes.enabled=true to get them everywhere, which also makes local testing of your probe configuration possible.

Do probes read the JSON body?

No. Kubernetes HTTP probes only look at the status code: 200 to 399 is success, anything else is failure. Spring Boot returns 200 for UP and 503 for DOWN or OUT_OF_SERVICE, so the body's "status" is mainly for humans. From a shell, curl -s -o /dev/null -w '%{http_code}' shows what the probe sees, and curl -sf exits non-zero on 4xx or 5xx.

What does timeoutSeconds do, and why is the default dangerous?

It is how long the kubelet waits for a probe response before counting a failure; the default is 1 second. A health endpoint that normally takes 50 ms but occasionally takes 1.2 s under load or during a GC pause will fail probes precisely when the service is busiest. Combined with a low failureThreshold that means restarts under load. Set 2-3 seconds and require a few consecutive failures for liveness.

In an interview Mid

What is the difference between liveness and readiness probes, and what should each check in a Spring Boot app?

The groups come from management.endpoint.health.probes.enabled=true (automatic when Boot detects Kubernetes). Probes read the status code: 200 UP, 503 DOWN.

The classic mistake: the aggregate /actuator/health (which includes db) as liveness. The database blips, every pod is restarted at once, they all come back into the same dead database. A restart does not fix a database.

Also: a startup probe for slow starters (failureThreshold x periodSeconds > start time), liveness tolerant of a few failures, the default timeoutSeconds: 1 in mind, and management.server.port so probes answer even when the request threads are stuck.

Also asked: Our database went down for a minute and every pod of the service restarted. Why, and how do you prevent it? · What is a startup probe for, and how do you size it for an app that takes 90 seconds to start? · Why would you run the management endpoints on a separate port?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.