Why this matters
Two everyday complaints: "stopping our container always takes ten seconds" and "the container I stopped for maintenance was running again after the reboot". Both are about signals and restart policies - things you already know from systemd, applied to containers.
What you need to know already: SIGTERM vs SIGKILL, and why PID 1 is special (3.6, 3.8); systemd's
Restart=andTimeoutStopSec(2.10, 2.24); exit codes and 128+N (11.1); ENTRYPOINT exec form vs shell form (10.28).
stop, kill, and what the process sees
docker stop api send the stop signal (SIGTERM), wait 10s, then SIGKILL
docker stop -t 30 api wait 30 seconds instead (-t = timeout)
docker kill api SIGKILL immediately
docker kill -s HUP web send any signal (-s): nginx reloads its config on HUP
The 10 seconds are the grace period: time for the app to finish what it is doing and exit by itself - the same idea as systemd's TimeoutStopSec (2.24).
An image can choose a different stop signal with STOPSIGNAL (10.28): nginx uses SIGQUIT (graceful: finish in-flight requests), postgres SIGINT (fast shutdown). docker inspect -f '{{.Config.StopSignal}}' web shows it.
What a stop means depends entirely on PID 1:
PID 1 handles SIGTERM exits cleanly: 0 or 143, in milliseconds
PID 1 has no handler the kernel IGNORES the signal for PID 1:
10 seconds, then SIGKILL, 137
PID 1 is a shell wrapping the app same as no handler: the app never hears it
A handler is code in the program that says "when SIGTERM arrives, do this". Normally a process without one dies on SIGTERM (the default action). PID 1 of a PID namespace is the exception (3.8): the kernel does not apply default actions to it, so without a handler the signal is simply ignored.
# node-app = a node image without a SIGTERM handler (not on this box)
time docker stop node-app
node-app
real 0m10.2s
docker inspect -f '{{.State.ExitCode}}' node-app
137
time (a bash keyword) measures how long the command took: real 10.2s = the full grace period, then SIGKILL, so exit 137.
--init
--init puts docker-init (it is tini, 3.8) in as PID 1. It forwards every signal to your process and reaps zombies. Your process becomes PID 2, where normal default actions apply - so a node or python app without a handler dies on SIGTERM instead of ignoring it:
docker run -d --init --name n node-app:1
docker top n
UID PID PPID C STIME TTY TIME CMD
1000 4410 4391 0 10:20 ? 00:00:00 /sbin/docker-init -- node server.js
1000 4422 4410 0 10:20 ? 00:00:00 node server.js
docker stop n ; docker inspect -f '{{.State.ExitCode}}' n
n
143
docker top columns are the ps -ef ones (3.1): the node process's PPID is docker-init's PID - tini is its parent. 143 (128 + 15), not 0: the process was killed by SIGTERM rather than exiting on purpose - fast, but no graceful shutdown ran. A real handler that closes the server and exits 0 is still the better fix; --init is the safety net.
Restart policies
A restart policy tells Docker what to do when a container's PID 1 exits - the container version of systemd's Restart= (2.10):
--restart no (default) never restart
--restart on-failure[:N] only after a non-zero exit, at most N times
--restart always after any exit - and after a daemon restart / reboot, even if you stopped it
--restart unless-stopped like always, but a manual docker stop is remembered
Restarts are not immediate. Docker waits 100ms before the first, and doubles the delay each time (200ms, 400ms, ... capped at one minute) - an exponential back-off, so a broken container does not burn the CPU restarting thousands of times. A container that then runs for at least 10 seconds resets the delay. While it waits, docker ps says so:
# app:bad = an image whose app exits 1 at startup (not on this box)
docker run -d --name flappy --restart always app:bad
docker ps
CONTAINER ID IMAGE COMMAND CREATED STATUS PORTS NAMES
2b3c4d5e6f7a app:bad "/app/server" 48 seconds ago Restarting (1) 6 seconds ago flappy
docker inspect -f '{{.RestartCount}}' flappy
7
Restarting (1) = the last exit code was 1; RestartCount = how many times Docker restarted it. A container stuck like this is in a crash loop. You can never exec into it (it is not running long enough) - read docker logs, which keeps the output of every run.
always vs unless-stopped, and why it matters on a reboot
The Docker daemon (dockerd, the background service that runs containers, 10.1) is a systemd unit. When it restarts - or the whole box reboots - it decides which containers to start again:
always unless-stopped
docker stop stays stopped stays stopped
daemon restart / reboot comes back stays stopped (you stopped it)
crash restarts restarts
always is how a container you stopped "for maintenance" is running again after Monday's kernel patch reboot. For long-running services on a single host, unless-stopped is usually what you mean.
restart and the "it works after a restart" trap
docker restart api is stop + start: same container, same writable layer, same config. It fixes nothing that is baked into the configuration - a wrong env var or a missing volume needs docker rm (delete the container) and a new docker run. That is called recreating the container. Restart vs recreate is also why Compose (11.31) compares the configuration and recreates a container when it changed.
What you can now do
- explain why a stop takes exactly 10 seconds, and fix it (handler,
exec,--init) - pick the right restart policy and predict what survives a reboot
- tell restarting a container from recreating it