OnCallReady

ProcessesLinuxDocker · 2 min read

Zombie processes (<defunct>): why kill doesn't work and what does

A zombie is a child that has exited but was never reaped by its parent. Why kill -9 does nothing, how to find the parent, and why containers need a PID 1 that reaps.

ps shows processes marked Z and <defunct>, and kill -9 on them does nothing. (Processes stuck in D ignore it too, for a different reason: see load 15 with an idle CPU.)

terminal
$ ps -eo pid,ppid,stat,comm | awk '$3 ~ /Z/'
 4410  4402 Z    worker <defunct>
 4411  4402 Z    worker <defunct>
 4412  4402 Z    worker <defunct>

What a zombie is

When a process exits, the kernel keeps a small record of it - its PID and exit status - until the parent reads that status with wait(). That reading is called reaping. Between the exit and the reaping, the dead child is a zombie. It uses no CPU and no memory, only an entry in the process table.

A few zombies that come and go are normal. Zombies that pile up mean the parent isn't reaping: a bug, or a parent stuck doing something else.

Why kill does nothing

You can't kill something that has already exited. Signals go to running processes, and a zombie has nothing left to receive them. kill -9 4410 "succeeds" and the line stays.

What works

Deal with the parent (the PPID column, 4402 above):

  • fix it or restart it, so it reaps its children;
  • or end it. Its children are then adopted by PID 1 (systemd, or a "subreaper"), and PID 1 reaps orphans as one of its jobs. The zombies vanish at once.
terminal
$ ps -o pid,comm -p 4402
  PID COMMAND
 4402 buggy-parent
$ kill 4402

The danger isn't the zombies' memory. It's running out of PIDs (pid_max, or a container's pids.max) when thousands pile up: then nothing in the box can fork, and a container can end up in CrashLoopBackOff.

Containers: PID 1 has a second job

Inside a container, your app is usually PID 1. PID 1 is expected to reap orphans and to forward signals, and most apps do neither. A container that spawns shell commands then fills up with zombies, and ignores SIGTERM on top of that, so every pod deletion waits out the full grace period. The fix is a tiny init as PID 1: tini (docker run --init adds it), or exec your app from the entrypoint script so it gets the signals directly.

OnCallReady is free, with no ads and no tracking. RSS · All posts