Why this matters
"The process is running" is not the same as "the app works". A web server can be up and stuck. A health check is a small command that asks the app "are you OK?"; Docker runs it for you and shows the answer. The golden signals from Ch 0 tell you users are hurting; a health check tells you which instance is.
What you need to know already: exit codes and || exit 1 (Ch 6), curl and HTTP status codes (Ch 9), base images without curl or a shell (previous two lessons), shell form vs exec form (the ENTRYPOINT lesson).
What it does
HEALTHCHECK tells Docker how to check the container. Docker runs the command inside the container on an interval and records the result:
HEALTHCHECK --interval=10s --timeout=3s --start-period=20s --retries=3 \
CMD curl -fsS http://localhost:8080/actuator/health || exit 1
--interval time between checks (default 30s)
--timeout a check taking longer fails (default 30s)
--start-period failures during it do not count (default 0s)
--retries consecutive failures to become unhealthy (default 3)
(curl's -f fails on HTTP errors like 500, -s is silent, -S still shows errors.) Exit 0 = healthy, exit 1 = unhealthy. || exit 1 turns curl's many exit codes into 1; exit code 2 is reserved. The state shows up in docker ps:
STATUS
Up 5 seconds (health: starting)
Up 2 minutes (healthy)
Up 9 minutes (unhealthy)
and the last five results, with their output, in docker inspect:
# api = the unhealthy container of the mission below
docker inspect -f '{{json .State.Health}}' api
{"Status":"unhealthy","FailingStreak":4,"Log":[{"Start":"2026-09-23T10:12:01.001Z","End":"2026-09-23T10:12:01.041Z","ExitCode":127,"Output":"/bin/sh: 1: curl: not found\n"}, ...]}
FailingStreak counts failures in a row; ExitCode 127 is "command not found"; Output is what the check printed. Read it before guessing. (Pipe it to jq, Ch 7, to make it readable.)
The trap: the check needs its tools
CMD curl ... runs curl in the image. eclipse-temurin ships curl. node:22-slim and python:3.12-slim do not. Alpine has busybox wget but no curl. Distroless has no shell at all, so the shell form cannot even start:
"Output": "OCI runtime exec failed: exec failed: unable to start container process: exec: \"/bin/sh\": stat /bin/sh: no such file or directory: unknown"
and the container is (unhealthy) forever while the application is perfectly fine. Options, in order of preference:
# alpine: busybox wget is there
HEALTHCHECK CMD wget -qO- http://localhost:8080/health || exit 1
# exec form, no shell needed, with a binary that exists in the image
HEALTHCHECK CMD ["/app/healthcheck"]
# slim images: install curl (costs a few MB) - or check with the runtime itself
HEALTHCHECK CMD ["node","-e","require('http').get('http://localhost:3000/health',r=>process.exit(r.statusCode===200?0:1)).on('error',()=>process.exit(1))"]
(wget -qO-: quiet, write the page to stdout.) The exec form (CMD ["..."]) runs without /bin/sh; the string form is called CMD-SHELL and needs one.
What health does NOT do
Plain Docker does not restart an unhealthy container. A restart policy (docker run --restart, like systemd's Restart=) reacts to the process exiting, not to health. Health is information: docker ps shows it, docker events (Docker's live event stream) emits health_status: unhealthy, and other tools can wait for it or act on it.
Later (Ch 11): Compose can wait until one container is healthy before it starts the next one that needs it.
The design lesson that carries everywhere: a health check must be cheap and local. Never make it call the database: then one slow database marks every copy of the app unhealthy at once, and a slowdown becomes an outage.
What you can now do
- Write a HEALTHCHECK with sensible timings.
- Read
.State.Healthto see why a check fails. - Pick a check that works in slim, alpine and distroless images.