/actuator/health
The problem. Lesson 17.22 said "liveness must be local, readiness may check dependencies". In a Spring Boot app that choice is a few properties - and the defaults change depending on whether the app knows it runs in Kubernetes.
What you need to know already: Actuator (21.1), the three probes and the restart storm (17.20, 17.22), HTTP status codes (9.21).
$ curl -s localhost:8080/actuator/health
{"status":"UP"}
By default you only see the aggregate. With details:
management.endpoint.health.show-details=always # or when-authorized
$ curl -s localhost:8080/actuator/health | jq .
{
"status": "UP",
"components": {
"db": { "status": "UP", "details": { "database": "PostgreSQL", "validationQuery": "isValid()" } },
"diskSpace": { "status": "UP", "details": { "total": 19338280960, "free": 11524853760, "threshold": 10485760, ... } },
"livenessState": { "status": "UP" },
"ping": { "status": "UP" },
"readinessState": { "status": "UP" },
"ssl": { "status": "UP", ... }
},
"groups": [ "liveness", "readiness" ]
}
Every HealthIndicator (a small built-in check) on the classpath (the set of libraries the app was built with - like its node_modules) contributes a component: db (a database connection, a "DataSource", is configured), diskSpace, redis, rabbit, kafka (caches and message queues)... The aggregate is DOWN if any component is DOWN - which is exactly why the aggregate must not be your liveness probe.
The HTTP status code carries the result: 200 for UP, 503 for DOWN or OUT_OF_SERVICE. Probes look at the status code, not the body.
The two groups
management.endpoint.health.probes.enabled=true
gives you two health groups (a group = a named subset of the components, with its own URL):
/actuator/health/liveness livenessState: CORRECT | BROKEN -> UP | DOWN
/actuator/health/readiness readinessState: ACCEPTING_TRAFFIC | REFUSING_TRAFFIC -> UP | OUT_OF_SERVICE
Boot enables them automatically when it detects it is running on Kubernetes (it sees the KUBERNETES_SERVICE_HOST variables). On a VM, in a test, or in Docker locally they are off until you set the property - and the paths 404:
# before probes.enabled is set (the actuator mission set it: yours answers UP now)
curl -s localhost:8080/actuator/health/liveness
{"timestamp":"...","status":404,"error":"Not Found","path":"/actuator/health/liveness"}
The states are driven by the application lifecycle: readiness becomes ACCEPTING_TRAFFIC only once the app has started (after its startup tasks, "runners"), and flips to REFUSING_TRAFFIC the moment graceful shutdown begins. Liveness goes BROKEN if the application publishes it (a fatal internal error).
Liveness vs readiness: the rule
liveness "restart me" must be cheap, LOCAL, and only fail if a restart would fix it
readiness "stop sending me traffic" may reflect dependencies the app cannot work without
- Never put a dependency in liveness. If the database blips and liveness includes
db, Kubernetes restarts every pod at once - and they all come back into the same dead database, hammering it. A restart does not fix a database. - Readiness may include critical dependencies:
management.endpoint.health.group.readiness.include=readinessState,db
management.endpoint.health.group.liveness.include=livenessState # the default; keep it
But think about it: if every pod goes unready at once, the Service has no endpoints and callers get connection errors instead of a fast 503 from you. Many teams keep readiness local too and handle dependency failure in code (circuit breaker, fallback).
- A health check that waits for a database connection can hang when the pool is exhausted. A probe with a 1-second timeout then fails - "the health check took the service down" is an incident in this chapter.
Running management on a separate port
management.server.port=8081
Separate port, separate (small) thread pool (a fixed set of threads that take turns doing the work, 20.20): a probe still answers when all 200 Tomcat request threads are stuck. With a separate port the probe paths also exist on the main port if you set management.endpoint.health.probes.add-additional-paths=true (/livez and /readyz).
The three probes for an app that takes 90 seconds to start
startupProbe:
httpGet: { path: /actuator/health/liveness, port: management }
periodSeconds: 5
failureThreshold: 30 # 30 x 5 s = 150 s allowed to start: 90 s plus margin
livenessProbe:
httpGet: { path: /actuator/health/liveness, port: management }
periodSeconds: 10
timeoutSeconds: 2
failureThreshold: 3 # 30 s of consecutive failure before a restart
readinessProbe:
httpGet: { path: /actuator/health/readiness, port: management }
periodSeconds: 5
timeoutSeconds: 2
failureThreshold: 2
- The startup probe holds off the other two until it succeeds once. Without it you either set
initialDelaySeconds: 100on liveness - which delays detecting a real hang forever after - or liveness kills the pod mid-startup and it never comes up (CrashLoopBackOff for a healthy app). - Liveness tolerates a GC pause or a burst: a few consecutive failures, not one.
- Readiness reacts fast, because taking a pod out of rotation is cheap.
- Timeouts matter: the kubelet's default
timeoutSecondsis 1. A health endpoint that sometimes takes 1.2 s under load fails its probe under load - precisely when you least want a restart.
Reading the status code from the shell
$ curl -s -o /dev/null -w '%{http_code}\n' localhost:8080/actuator/health/readiness
200
$ curl -sf localhost:8080/actuator/health/liveness >/dev/null && echo alive || echo dead
alive
-f makes curl exit non-zero on 4xx/5xx - that is what a shell probe (or a systemd watchdog script) relies on. -o /dev/null -w '%{http_code}\n' = throw the body away and print only the status code.
What you can now do
- Turn on the liveness and readiness groups and choose what each includes.
- Put management on its own port so probes survive a busy app.
- Write the three probes for a slow-starting Spring Boot app and justify every number.