OnCallReady

Lesson 11.1 · Docker Runtime & Networking · 19 min read

Running containers, and reading how they ended

In plain words

Think of a wind-up toy. You wind it, put it down, and it walks until the spring runs out. When it stops, it stays on the table where it stopped. It does not vanish, you just stop noticing it among the moving ones.

A container is the same. docker run creates it and starts its main process, PID 1. The container runs exactly as long as PID 1 does, and when PID 1 ends, the container stops and keeps the exit code as a note of how it ended. docker ps only shows the toys still walking; docker ps -a shows the ones that stopped, with Exited (137) or Exited (1) in the STATUS column telling you why.

Why this matters

Chapter 10 was about building images. This chapter is about running them - and about the first ticket everyone gets: "I started my container and it disappeared". It did not disappear. It stopped, and it left a note saying why. This lesson teaches you where that note is and how to read it.

What you need to know already: a container is an ordinary Linux process started from an image, fenced in by namespaces and cgroups (10.3); docker run and the difference between ENTRYPOINT and CMD (10.28); exit statuses and $? (6.5); signals like SIGTERM and SIGKILL (3.6).

The flags you will actually type

docker run IMAGE [COMMAND] makes a new container from an image and starts it. These are the flags you will see all chapter - skim them now, each one gets its own lesson later:

-d                 detached: run in the background, print the container ID
--rm               delete the container as soon as it exits
--name web         give it a stable name instead of a random one (jolly_hopper)
-p 8080:80         publish a port: HOST_PORT:CONTAINER_PORT (lesson 11.15)
-e KEY=value       set an environment variable      --env-file app.env   many at once
-v name:/path      mount a named volume             -v /host/dir:/path   a bind mount (11.22)
--network appnet   which Docker network to join (11.15)
-u 1000:1000       run as this uid:gid (user id : group id, as in 4.3)
-m 512m --cpus 1   memory and CPU limits - cgroups again (11.10)
--restart unless-stopped   restart policy: bring it back when it dies (11.7)
--entrypoint sh    replace the image's ENTRYPOINT with another program
--init             run a tiny init (tini) as PID 1 (3.8, 11.7)
-it                -i keep stdin open + -t give it a terminal: needed for a shell

Detached mode means the container runs in the background and your prompt comes back straight away; without -d your terminal is attached to the container's output until it ends.

docker run is really two steps: docker create (make the container object from the image, with all your flags) plus docker start (start its main process). The container object exists from the create onwards - which is why a container that failed to start still shows up in the list below.

The lifecycle

Every container is in one container state at a time:

created  ->  running  ->  exited
                ^  |
                |  v
             restarting        (a restart policy is bringing it back)

docker ps lists containers - but only running and restarting ones. docker ps -a (-a = all) includes the stopped ones. Forgetting -a is the single most common reason people think "my container disappeared":

$ docker run -d --name batch reports:3 python export.py
$ docker ps
CONTAINER ID   IMAGE   COMMAND   CREATED   STATUS   PORTS   NAMES
$ docker ps -a
CONTAINER ID   IMAGE       COMMAND                  CREATED          STATUS                      PORTS   NAMES
4f1e2d3c4b5a   reports:3   "python export.py"       20 seconds ago   Exited (1) 18 seconds ago           batch

The columns: CONTAINER ID (the first 12 characters of its ID), IMAGE it was made from, COMMAND it ran as PID 1, CREATED how long ago, STATUS (state, exit code, age), PORTS it publishes, NAMES.

STATUS is your first diagnosis:

A container lives exactly as long as PID 1

Inside its PID namespace, the container's main process is PID 1 (10.3, 10.28). There is no such thing as "the container is up but the app crashed": when PID 1 exits, the container exits, and it keeps PID 1's exit status.

$ docker run ubuntu:24.04
$ docker ps -a --filter ancestor=ubuntu:24.04
CONTAINER ID   IMAGE          COMMAND       CREATED         STATUS                     PORTS   NAMES
8a7b6c5d4e3f   ubuntu:24.04   "/bin/bash"   3 seconds ago   Exited (0) 2 seconds ago           keen_hopper

(--filter ancestor=IMAGE shows only containers made from that image.)

The image's CMD is /bin/bash. bash started with no terminal and no script, read "end of input" from stdin straight away, and exited 0 - correct behaviour. docker run -it ubuntu:24.04 gives it a terminal, and you get a shell.

The same trap catches programs that daemonise: they fork a copy of themselves into the background and let the original process exit (the classic way to start a server before systemd). In a container the original process is PID 1, so the container ends at once. That is why the nginx image runs nginx -g 'daemon off;': stay in the foreground.

Exit codes: the table to know by heart

You met exit statuses in bash (6.5) and in systemctl status (2.9). A container's exit code is simply PID 1's exit status:

0      finished normally
1      the application failed (generic) - read the logs
2      usage error: bad flags or config (many CLIs, shell builtins)
125    docker itself failed: bad flag, name already taken, port in use
126    the command exists but cannot be executed (no execute bit, or it is a directory)
127    the command was not found in the image
128+N  killed by signal number N:
         130  SIGINT  (2, Ctrl+C)
         134  SIGABRT (6, the program aborted itself: node's "heap out of memory")
         137  SIGKILL (9, docker kill, the stop timeout, or the OOM killer)
         139  SIGSEGV (11, a segfault, often in a native library)
         143  SIGTERM (15, a normal docker stop)
255    the runtime could not start the entrypoint (exec format error, missing interpreter)

128+N is the shell convention from 6.5: when a signal kills a process, its status is 128 plus the signal number (kill -l in 3.6 lists the numbers). A segfault (SIGSEGV) is the kernel killing a program that touched memory it does not own - a bug in the program, usually in native (C/C++) code.

137 is ambiguous on purpose: killed by SIGKILL, which is either a person (docker kill, or docker stop giving up) or the kernel's OOM killer (5.7, 5.11). docker inspect prints everything Docker knows about a container as JSON; -f (format) picks out single fields. Here, how it ended:

# api = an OOM-killed container (the limits mission later)
docker inspect -f '{{.State.ExitCode}} {{.State.OOMKilled}} {{.State.Error}}' api
137 true

The {{ }} syntax is a Go template; lesson 11.5 covers it.

The failure modes, with the output that identifies them

$ docker run --name t1 alpine:3.20 sever
docker: Error response from daemon: failed to create task for container: failed to create shim task: OCI runtime create failed: runc create failed: unable to start container process: error during container init: exec: "sever": executable file not found in $PATH: unknown.
$ echo $?
127
$ docker ps -a --filter name=t1 --format '{{.Status}}'
Created

Read long Docker errors from the end: exec: "sever": executable file not found in $PATH - a typo for server. (runc is the low-level program that actually creates the namespaces and starts the process; you will see its name in errors.) The container never left Created. The same shape with permission denied is exit 126 (the file is not executable). With no such file or directory for a script that exists, the interpreter in its shebang line is missing (10.28).

$ docker run -d --name t2 postgres:16
$ docker logs t2
Error: Database is uninitialized and superuser password is not specified.
       You must specify POSTGRES_PASSWORD to a non-empty value for the
       superuser. For example, "-e POSTGRES_PASSWORD=password" on "docker run".

(Postgres is a popular open-source SQL database; its image refuses to start without a password.) This time the application ran and chose to exit 1. docker logs NAME prints what the container's PID 1 wrote to its output - and it works on exited containers. It is the first thing to run.

The diagnosis, in order

docker ps -a                                    it IS there - status and exit code
docker logs <c>                                 what it said before dying
docker inspect -f '{{.State.ExitCode}} {{.State.OOMKilled}} {{.State.Error}}' <c>
docker inspect -f '{{json .Config.Entrypoint}} {{json .Config.Cmd}}' <c>   what it ran
docker run --rm -it --entrypoint sh <image>     get in and look around

The question is never "why did it stop" but what was PID 1, and what did it do? Those five commands answer it.

What you can now do

Why it helps

The most common "Docker is broken" ticket is a container that "disappeared". Knowing that it is sitting in docker ps -a with an exit code turns a vague complaint into a diagnosis in one command.

The exit code table pays off every week: 127 means a typo in the command or a missing binary in the image, 126 a permission bit, 137 a SIGKILL that is either OOM or a stop timeout, 143 a clean SIGTERM, 1 the app itself giving up. The same numbers show up in systemctl status, in CI jobs that "fail for no reason", and in interview questions about exit 137. And "a container lives as long as PID 1" explains every image whose process daemonises and exits instantly.

Commands in this lesson

docker echo

FAQ

Why does my container exit immediately after docker run?

Because PID 1 finished. Maybe it was a one-shot command that completed, a shell with no terminal that read end-of-file and exited 0, or a server that forked into the background and let its parent exit, which ends the container. Check docker ps -a for the code and docker logs for the output. For shells use -it; for daemons run them in the foreground, like nginx -g 'daemon off;'.

What is the difference between exit 125, 126 and 127?

125 means Docker itself failed before your process ran: a bad flag, a name that is already taken, a port already allocated. 126 means the command was found but could not be executed, usually a missing execute bit or a path that is a directory. 127 means the command was not found in the image's PATH at all, often a typo or a tool that the base image does not contain. In the 126 and 127 cases the container stays in Created.

Is exit 137 always an out-of-memory kill?

No. 137 is 128 + 9, meaning the process received SIGKILL. That can be the cgroup OOM killer, but also docker kill, or docker stop running out of grace time because PID 1 ignored SIGTERM. docker inspect -f '{{.State.OOMKilled}}' separates them: true means the memory limit, false means someone or something sent the kill. docker events shows an oom event right before the die when it was memory.

What does docker run do that docker start does not?

docker run is create plus start plus attach: it builds a new container object from an image, with all the flags you give it, and then starts it. docker start restarts an existing, stopped container with the configuration it already has. That is why a failed run still leaves a container behind in docker ps -a, and why a new -e or -p needs a new container, not a start.

Should I always use --rm?

For throwaway runs, yes: one-off commands, debug shells, tests. The container is deleted when it exits, so your docker ps -a does not fill with dead ones. For long-running services, no: without the stopped container you lose docker logs and docker inspect after a crash, and those are exactly what you need to find out why it died. --rm also cannot be combined with a restart policy.

In an interview Junior

What does exit code 137 mean for a container, and how do you find out what caused it?

137 = 128 + 9: PID 1 was killed by SIGKILL. Two possible senders:

Tell them apart:

docker inspect -f '{{.State.ExitCode}} {{.State.OOMKilled}}' api

OOMKilled: true means memory. Then docker logs api for what it printed before.

The rest of the table: 0 finished normally, 1 the app failed (logs), 126 not executable, 127 command not found (the container never left Created), 139 segfault (128 + 11), 143 SIGTERM (128 + 15). And remember docker ps hides stopped containers: docker ps -a.

Also asked: A container is not in docker ps. Where did it go? · Why does a container started from the ubuntu image exit immediately? · What is the difference between docker create, docker start and docker run?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.