OnCallReady

Lesson 10.28 · Images & Builds · 21 min read

ENTRYPOINT, CMD, and who is PID 1

In plain words

Imagine a class where only the class captain hears the teacher's announcements, and the captain is supposed to pass them on. If the captain is a kid who never listens and never passes anything on, the rest of the class keeps playing when the bell rings, until the teacher comes in and drags everyone out.

In a container, PID 1 is the captain. docker stop sends SIGTERM only to PID 1, and the kernel ignores signals PID 1 has no handler for. With shell-form ENTRYPOINT java -jar app.jar, PID 1 is /bin/sh, which neither handles nor forwards SIGTERM, so Docker waits 10 seconds and SIGKILLs (exit 137). Exec form makes java PID 1; --init or tini adds a real captain; entrypoint scripts end with exec "$@".

Why this matters

"Every time we deploy, a few requests fail." "Stopping the container always takes exactly 10 seconds." Both are usually one Dockerfile line: the wrong process is PID 1, so it never hears the polite stop signal and gets killed mid-request. You already know the pieces - signals and PID 1 from Ch 3 - this lesson puts them inside a container.

What you need to know already: SIGTERM, SIGKILL and signal handlers (Ch 3, "Signals"), PID 1 and reaping zombies (Ch 3, "Zombies, orphans and PID 1"), exit code 137 = 128 + 9 (Ch 5, "cgroup OOM and exit 137"), exec and "$@" in bash (Ch 6), the PID namespace ("What a container actually is").

ENTRYPOINT is the program, CMD is its default arguments

The container runs ENTRYPOINT + CMD, joined together:

ENTRYPOINT ["java","-jar","app.jar"]
CMD ["--spring.profiles.active=prod"]
docker run app                       java -jar app.jar --spring.profiles.active=prod
docker run app --debug               java -jar app.jar --debug          (CMD replaced)
docker run --entrypoint sh app       sh                                 (ENTRYPOINT replaced, CMD cleared)
docker run -it --entrypoint sh app -c 'ls /app'    sh -c 'ls /app'

Anything after the image name replaces CMD, never ENTRYPOINT. --entrypoint replaces the ENTRYPOINT and throws away the image's CMD. (-it = -i keep input open + -t give it a terminal, so you can type into it.) docker inspect and docker ps show what was really started:

$ docker inspect -f '{{json .Config.Entrypoint}} {{json .Config.Cmd}}' web
["/docker-entrypoint.sh"] ["nginx","-g","daemon off;"]
$ docker ps --no-trunc --format '{{.Command}}'
"/docker-entrypoint.sh nginx -g daemon off;"

Only CMD (no ENTRYPOINT) is common for general-purpose images: ubuntu has CMD ["/bin/bash"], so docker run ubuntu ls runs ls instead.

Exec form vs shell form

ENTRYPOINT ["java","-jar","app.jar"]    # exec form: a JSON array
ENTRYPOINT java -jar app.jar            # shell form: a string

The exec form runs the program directly. The shell form is rewritten to ["/bin/sh","-c","java -jar app.jar"]. So the container's PID 1 is /bin/sh, and java is its child:

# shellform / execform: the two containers of the mission below
docker top shellform
UID    PID    PPID   C   STIME   TTY   TIME       CMD
root   3120   3101   0   10:02   ?     00:00:00   /bin/sh -c java -jar app.jar
root   3141   3120   0   10:02   ?     00:00:04   java -jar app.jar

(PID 3141's parent, PPID, is 3120: the shell.) Two things follow:

  1. sh does not forward signals. docker stop sends SIGTERM to PID 1 - the shell. The Java process never hears about it.
  2. PID 1 is special in a PID namespace. The kernel does not apply the default action of a signal (for SIGTERM: die) to PID 1. A signal PID 1 has no handler for (code that runs when the signal arrives) is simply ignored. sh installs no SIGTERM handler, so SIGTERM does nothing at all.

Docker waits the grace period (10 seconds by default; docker stop -t N or docker run --stop-timeout N to change it), then sends SIGKILL:

time docker stop shellform
shellform

real	0m10.2s
docker ps -a --filter name=shellform --format '{{.Status}}'
Exited (137) 3 seconds ago

(time measures how long a command takes; docker ps -a includes stopped containers; --filter name=... keeps only matching ones.) The exec form makes java PID 1. The JVM (the Java virtual machine, the program that runs Java code) installs a SIGTERM handler, runs its shutdown code (the app closes its web server and database connections, flushes logs) and exits:

time docker stop execform
execform

real	0m0.4s
docker ps -a --filter name=execform --format '{{.Status}}'
Exited (143) 2 seconds ago

Exit codes above 128 are "killed by signal N": 137 = 128 + 9 (SIGKILL), 143 = 128 + 15 (SIGTERM). A 137 after a stop means your process ignored SIGTERM and was killed with requests still in flight - it is why "every deploy drops connections". (The same number from Ch 5: a 137 without a stop is usually the OOM killer.)

BuildKit warns about it on every build: JSONArgsRecommended: JSON arguments recommended for ENTRYPOINT to prevent unintended behavior related to OS signals.

The PID 1 trap is not only about shells

Any program that is PID 1 and has no SIGTERM handler ignores docker stop:

java      installs handlers           stops on SIGTERM, 143
nginx     handles TERM/QUIT           image sets STOPSIGNAL SIGQUIT, exits 0
postgres  handles INT/TERM            image sets STOPSIGNAL SIGINT, exits 0
node      NO handler by default       ignored as PID 1 -> 10s -> 137
python    NO handler by default       ignored as PID 1 -> 10s -> 137
sh/bash   NO handler                  ignored as PID 1 -> 10s -> 137

(postgres is a database; node runs JavaScript; python runs Python.) For node, add the handler (process.on('SIGTERM', ...)) - which you want anyway, to close the server gracefully. For python, signal.signal(signal.SIGTERM, ...). It is the same idea as trap ... TERM in a bash script (Ch 6). Or put a real init in front.

tini: an init for one process

An init is the first process of a system - systemd on this box. tini is a tiny init made for containers. docker run --init puts docker-init (tini) in as PID 1. It forwards signals to your process and reaps zombies (the PID 1 duties from Ch 3). Your program is now PID 2, where default signal actions apply again - so even a handler-less node dies on SIGTERM:

# node-app:1 = a node image without a SIGTERM handler (not on this box)
docker run -d --init --name n node-app:1
docker top n
UID    PID    PPID   C   CMD
1000   3310   3291   0   /sbin/docker-init -- node server.js
1000   3322   3310   0   node server.js
docker stop n && docker ps -a --filter name=n --format '{{.Status}}'
Exited (143) 1 second ago

In a Dockerfile: RUN apt-get install -y tini (or apk add tini on alpine) and ENTRYPOINT ["/usr/bin/tini","--","node","server.js"]. --init is a flag of docker run, and not every place that runs images has it, so if you rely on an init, bake it into the image.

Entrypoint scripts: always end with exec

Many images need setup before the app starts - filling in a config file, waiting for something, fixing permissions. The pattern is a small shell script as the ENTRYPOINT:

#!/bin/sh
set -e
envsubst < /etc/app/app.conf.tpl > /etc/app/app.conf
exec "$@"

(envsubst replaces $VARS in a file with their values.)

COPY --chmod=755 docker-entrypoint.sh /usr/local/bin/
ENTRYPOINT ["docker-entrypoint.sh"]
CMD ["java","-jar","/app/app.jar"]

exec replaces the shell with the program (same PID), so java becomes PID 1 and gets the signals. "$@" is the script's arguments - here, the CMD - which is why this combination is so flexible: docker run app bash runs the setup and then bash. The official images all do this - /docker-entrypoint.sh in nginx, docker-entrypoint.sh in postgres and node, /__cacert_entrypoint.sh in eclipse-temurin all end in exec "$@".

Forget the exec (last line java -jar /app/app.jar) and you are back to the shell-form bug: sh stays PID 1 and eats SIGTERM.

Scripts that will not start

exec /usr/local/bin/docker-entrypoint.sh: no such file or directory

The file is there - docker run --entrypoint ls app -l /usr/local/bin shows it. What is missing is the interpreter named in its first line (the shebang, Ch 6). Three causes:

And exec: "/app/start.sh": permission denied (exit 126) means the file is not executable (Ch 4): COPY --chmod=755, or chmod +x it in git.

What you can now do

Why it helps

This is behind a very common production complaint: 'every deploy drops a few requests' or 'stopping the app always takes exactly 10 seconds'. The symptom is exit code 137 after a normal stop, and the cause is PID 1 not handling SIGTERM, so Docker waits out the grace period and SIGKILLs the requests in flight. With this lesson you check docker top, see /bin/sh -c as PID 1, and fix the Dockerfile with one line. You will also debug the startup errors by sight: exec ... no such file or directory for a script that exists (missing interpreter, CRLF line endings, missing libc) and exit 126 for a missing execute bit. And you understand ENTRYPOINT versus CMD well enough to run docker run --entrypoint sh to get into a broken image.

Commands in this lesson

docker

FAQ

What is the difference between ENTRYPOINT and CMD?

ENTRYPOINT is the program; CMD is its default arguments, and the container runs the two concatenated. Anything after the image name in docker run replaces CMD, never ENTRYPOINT. --entrypoint replaces the ENTRYPOINT and also clears the image's CMD. With only CMD set, as in ubuntu, arguments after the image name replace the whole command.

Why does exit code 137 or 143 show up after docker stop?

Codes above 128 mean killed by signal N, where N is the code minus 128. 143 is SIGTERM (15): the process received the stop signal and exited. 137 is SIGKILL (9): the process ignored SIGTERM for the whole grace period and was killed. A 137 after a stop is the signal of a PID 1 problem; a 137 without a stop is usually the OOM killer.

Why does node or python ignore docker stop when java does not?

The kernel does not apply a signal's default action to PID 1 in a PID namespace; a signal without a handler is ignored. The JVM installs a SIGTERM handler for its shutdown hooks, so it exits. Node and Python install none by default, and neither does sh. Add a handler (process.on('SIGTERM'), signal.signal) or run behind an init like tini.

What does exec "$@" at the end of an entrypoint script do?

exec replaces the shell process with the program, keeping the same PID, so the app becomes PID 1 and receives signals directly. "$@" is the arguments passed to the script, which is the CMD. So the script runs setup, then turns into whatever CMD says. Without exec, sh stays PID 1 and swallows SIGTERM, the same bug as shell form.

Can I rely on docker run --init everywhere my image runs?

No. --init is a docker run flag that injects docker-init (tini) as PID 1, and other tools that start containers may not offer it. If your process needs an init, bake it into the image: install tini and use ENTRYPOINT ["/usr/bin/tini", "--", "node", "server.js"]. Better still, make the app handle SIGTERM itself, which it needs for graceful shutdown anyway.

In an interview Junior

What is the difference between the exec form and the shell form of ENTRYPOINT, and why does it matter?

docker stop sends SIGTERM to PID 1. sh does not forward it, and PID 1 in a PID namespace ignores signals it has no handler for. So nothing happens for the 10-second grace period, then SIGKILL: exit 137 (128 + 9), requests killed mid-flight. With the exec form the JVM handles SIGTERM, shuts down cleanly and exits 143.

Fixes: exec form; an entrypoint script that ends in exec "$@"; or an init like tini (docker run --init) for programs without a SIGTERM handler. BuildKit warns with JSONArgsRecommended.

And: ENTRYPOINT is the program, CMD its default arguments; arguments after the image name replace CMD.

Also asked: What is the difference between ENTRYPOINT and CMD? · What do exit codes 137 and 143 mean for a container? · Why should an entrypoint script end with exec "$@"?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.