OnCallReady

Lesson 10.44 · Images & Builds · 10 min read

HEALTHCHECK

In plain words

Imagine a night nurse who pokes her head into a patient's room every ten seconds and asks "are you OK?". If the patient answers, she writes "fine" on the chart. If three checks in a row get no answer, she writes "unwell". She does not treat the patient herself: she only keeps the chart, and it is up to the doctor to act on it. And she asks in the patient's own language, so if the room has no telephone, she cannot call.

HEALTHCHECK is that nurse. Docker runs your check command inside the container on an interval, records healthy or unhealthy after --retries failures, and shows it in docker ps and docker inspect. Plain Docker does not restart anything because of it, and the check can only use tools that exist in the image.

Why this matters

"The process is running" is not the same as "the app works". A web server can be up and stuck. A health check is a small command that asks the app "are you OK?"; Docker runs it for you and shows the answer. The golden signals from Ch 0 tell you users are hurting; a health check tells you which instance is.

What you need to know already: exit codes and || exit 1 (Ch 6), curl and HTTP status codes (Ch 9), base images without curl or a shell (previous two lessons), shell form vs exec form (the ENTRYPOINT lesson).

What it does

HEALTHCHECK tells Docker how to check the container. Docker runs the command inside the container on an interval and records the result:

HEALTHCHECK --interval=10s --timeout=3s --start-period=20s --retries=3 \
  CMD curl -fsS http://localhost:8080/actuator/health || exit 1
--interval      time between checks            (default 30s)
--timeout       a check taking longer fails    (default 30s)
--start-period  failures during it do not count (default 0s)
--retries       consecutive failures to become unhealthy (default 3)

(curl's -f fails on HTTP errors like 500, -s is silent, -S still shows errors.) Exit 0 = healthy, exit 1 = unhealthy. || exit 1 turns curl's many exit codes into 1; exit code 2 is reserved. The state shows up in docker ps:

STATUS
Up 5 seconds (health: starting)
Up 2 minutes (healthy)
Up 9 minutes (unhealthy)

and the last five results, with their output, in docker inspect:

# api = the unhealthy container of the mission below
docker inspect -f '{{json .State.Health}}' api
{"Status":"unhealthy","FailingStreak":4,"Log":[{"Start":"2026-09-23T10:12:01.001Z","End":"2026-09-23T10:12:01.041Z","ExitCode":127,"Output":"/bin/sh: 1: curl: not found\n"}, ...]}

FailingStreak counts failures in a row; ExitCode 127 is "command not found"; Output is what the check printed. Read it before guessing. (Pipe it to jq, Ch 7, to make it readable.)

The trap: the check needs its tools

CMD curl ... runs curl in the image. eclipse-temurin ships curl. node:22-slim and python:3.12-slim do not. Alpine has busybox wget but no curl. Distroless has no shell at all, so the shell form cannot even start:

"Output": "OCI runtime exec failed: exec failed: unable to start container process: exec: \"/bin/sh\": stat /bin/sh: no such file or directory: unknown"

and the container is (unhealthy) forever while the application is perfectly fine. Options, in order of preference:

# alpine: busybox wget is there
HEALTHCHECK CMD wget -qO- http://localhost:8080/health || exit 1

# exec form, no shell needed, with a binary that exists in the image
HEALTHCHECK CMD ["/app/healthcheck"]

# slim images: install curl (costs a few MB) - or check with the runtime itself
HEALTHCHECK CMD ["node","-e","require('http').get('http://localhost:3000/health',r=>process.exit(r.statusCode===200?0:1)).on('error',()=>process.exit(1))"]

(wget -qO-: quiet, write the page to stdout.) The exec form (CMD ["..."]) runs without /bin/sh; the string form is called CMD-SHELL and needs one.

What health does NOT do

Plain Docker does not restart an unhealthy container. A restart policy (docker run --restart, like systemd's Restart=) reacts to the process exiting, not to health. Health is information: docker ps shows it, docker events (Docker's live event stream) emits health_status: unhealthy, and other tools can wait for it or act on it.

Later (Ch 11): Compose can wait until one container is healthy before it starts the next one that needs it.

The design lesson that carries everywhere: a health check must be cheap and local. Never make it call the database: then one slow database marks every copy of the app unhealthy at once, and a slowdown becomes an outage.

What you can now do

Why it helps

The classic incident: a container shows (unhealthy) forever, someone spends an hour on the app, and docker inspect -f '{{json .State.Health}}' shows curl: not found because the base image was switched to slim or distroless. Knowing that the check runs inside the image saves you that hour. Health also feeds other tools: they can wait for a database to be healthy before starting the app that needs it, so a broken check can mean a stack that never starts. And the design rule you learn here applies to every health check you will ever write: it must be cheap and local. A check that calls the database is how one slow dependency marks every copy of the app unhealthy at once, and a slowdown becomes an outage.

FAQ

Does Docker restart a container when it becomes unhealthy?

No. --restart policies react to the main process exiting, not to health status, just like systemd's Restart=. Health is information: docker ps shows it, docker events emits health_status: unhealthy, and other tools can wait for it or act on it. If you need a restart when the app is unhealthy, something outside plain Docker has to watch the status and do it.

Should my health check also test the database?

Usually not. A health check should answer one question: is this copy of the app broken? If it also checks the database, then one slow database makes every copy report unhealthy at the same moment, and anything acting on health treats a slowdown as a total failure. Keep it cheap and local, like a small /health URL, and watch dependencies separately with the golden signals from Ch 0.

Why is my container unhealthy when the app works fine?

Read the check's own output first: docker inspect -f '{{json .State.Health}}' api | jq shows the last five results with exit codes and output. The usual causes are a missing tool (exit 127, curl: not found), a shell-form check in an image with no /bin/sh (distroless), the wrong port or path, or an app that listens only on a different interface. Also check --start-period for slow JVM starts.

What is the difference between the exec form and the shell form of HEALTHCHECK CMD?

HEALTHCHECK CMD curl -f http://localhost/ || exit 1 is the shell form: Docker runs it with /bin/sh -c, which is what makes || exit 1 work, and it fails in images with no shell. HEALTHCHECK CMD ["/app/healthcheck"] is the exec form: it runs the binary directly, with no shell features. For distroless, ship a small static health binary or use the runtime itself in exec form.

What do the timing options actually mean?

--interval is the time between checks (default 30s). --timeout fails a check that takes longer (default 30s). --retries is how many consecutive failures turn the status to unhealthy (default 3). --start-period is a grace time after start during which failures do not count toward the retries (default 0s), though a success ends it early. For a JVM that takes 20s to start, set a start period, or it may be marked unhealthy before it ever served a request.

In an interview Junior

What does HEALTHCHECK do in a Dockerfile, and what are its pitfalls?

It tells Docker to run a command inside the container on an interval and record the result: exit 0 healthy, exit 1 unhealthy.

HEALTHCHECK --interval=10s --timeout=3s --start-period=20s --retries=3 \
  CMD curl -fsS http://localhost:8080/actuator/health || exit 1

--start-period gives a slow starter time before failures count. The state shows in docker ps (health: starting, healthy, unhealthy); docker inspect -f '{{json .State.Health}}' shows the last results with their output.

Pitfalls:

Also asked: How would you set the HEALTHCHECK options for a service that takes 20 seconds to start? · Does Docker restart a container that becomes unhealthy? · Why should a health check not depend on the database?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.