OnCallReady

Lesson 3.8 · Processes & Signals · 6 min read

Zombies, orphans and PID 1

In plain words

When someone leaves a club, the club keeps a tiny card with "left, and here is how they did" until the person who invited them comes to read it. The member is gone, uses no chair and no drinks, but the card stays in the drawer. Shouting at the card does nothing. Only the inviter reading it, or the club manager taking over the inviter's cards, clears it.

That card is a zombie, state Z in ps: an exited process whose parent has not called wait(). kill -9 on it does nothing, because it is already dead. Fix or kill the parent and the zombies are handed to PID 1, the club manager, who reads every card immediately. Orphans are the opposite: the inviter left first, so the member is adopted by PID 1.

The process that is dead but will not go away

Sooner or later ps shows you lines marked <defunct> that no kill can remove. They are zombies, and the fix is never to shoot harder - it is to look at their parent.

What you need to know already: 1.15 (PID 1, parents), 3.1 (STAT Z), 3.6 (signals).

A zombie is not a running process

When a process exits, the kernel frees its memory, its open files, everything - but keeps one small record: the exit status (the number it ended with, 1.7). That record exists so the parent can call wait() - a system call meaning "tell me how my child ended" - and find out.

Until the parent does that, the entry stays in the process table as Z, <defunct>. It uses no CPU and no memory. It costs exactly one PID. Collecting the exit status is called reaping.

So:

A handful of zombies is cosmetic. Thousands mean a parent that will eventually use up every PID the kernel can hand out, and then nothing on the box can start a new process.

Orphans

If a parent exits first, its still-running children are orphans, and the kernel immediately makes PID 1 their new parent (re-parenting). You can see it in ps -ef: a process with PPID 1 that is obviously not a system service was orphaned, usually because someone's SSH session dropped or a wrapper script exited early.

PID 1's two jobs

  1. Reap. Anything re-parented to it gets wait()ed on, so orphans never linger as zombies.
  2. Pass signals on at shutdown: when the machine stops, it tells every service to stop (systemd sends SIGTERM, waits, then SIGKILL - 2.24).

systemd does both. A normal application does neither - which only matters when an application ends up being PID 1.

Later (Ch 10): that is exactly what happens inside a container: the program you start becomes PID 1 of its own small process table. A shell script as PID 1 does not pass SIGTERM on to the program it started (so a stop waits the full grace period and ends in SIGKILL), and an application as PID 1 rarely reaps the orphans it inherits (container zombies). The fixes - exec the real program, or a tiny init such as tini / docker run --init - are covered with Docker.

PID 1 is immune

sudo kill -9 1

Nothing happens. The kernel does not deliver signals to PID 1 unless PID 1 has explicitly installed a handler for them - a deliberate safeguard, because if PID 1 dies the kernel stops the whole machine (a kernel panic).

Your own shell is protected differently: an interactive bash ignores SIGTERM, so kill $$ does nothing. kill -9 $$ cannot be ignored, and it ends your session.

What you can now do

Why it helps

Zombies look alarming in ps, and the instinct is to kill them, which does nothing. Recognising that a few are harmless stops wasted effort, and knowing that the fix is always the parent turns a confusing screen into a one-command fix.

The real danger is the growing kind: a service that forks helpers and never reaps them can pile up thousands of zombies until the PID limit is hit and nothing on the box can start a process, which shows up as "fork: Resource temporarily unavailable". Knowing PID 1's duties (reap, pass signals on) also explains why init is protected, and the kill $$ versus kill -9 $$ difference explains how you can lose your own session by accident.

Commands in this lesson

kill

FAQ

Do zombies use memory or CPU?

Almost none. When a process exits, the kernel frees its memory, open files and most kernel structures. What remains is a small entry with the PID and exit status for the parent to collect. The real cost is the PID: each zombie holds one, and a system or a unit with a PID limit (kernel.pid_max, systemd's TasksMax=) eventually cannot create new processes if zombies pile up.

How do I find the parent of a zombie?

ps -o pid,ppid,stat,comm -p ZPID shows its PPID, or list all of them with ps -eo pid,ppid,stat,comm | awk '$3 ~ /Z/'. pstree -p shows zombies under their parent as {name} or <defunct>. The parent is the process that should be calling wait(); restarting or fixing it clears its zombies.

Can I force the parent to reap without killing it?

Sometimes. Sending the parent SIGCHLD (kill -CHLD PPID) can nudge a program whose handler reaps on that signal but missed one. If the program simply never calls wait(), nothing external can make it. Then the options are restarting the parent or, as a last resort, debugging it with gdb to call waitpid manually, which is rarely worth it.

What is a subreaper?

A process that has asked the kernel (with the prctl(PR_SET_CHILD_SUBREAPER) system call) to adopt its orphaned descendants instead of PID 1. systemd's per-user manager (systemd --user) and some supervisors use it, so orphans stay inside their own group and get reaped there. It is why an orphan's new PPID is not always 1: it is the nearest subreaper above it.

Why does kill $$ not end my shell but kill -9 $$ does?

An interactive bash ignores SIGTERM, so an accidental kill of your own shell does nothing. SIGKILL cannot be ignored, so kill -9 $$ ends the shell immediately, closing your SSH session. Because bash dies before it can forward SIGHUP to its jobs, background jobs become orphans re-parented to PID 1, and may die later when they write to the vanished terminal.

In an interview Junior

What is a zombie process, and how do you get rid of it?

A zombie (Z, <defunct> in ps) is a process that has already exited: the kernel freed its memory and files but keeps one small record, the exit status, until the parent collects it with wait() - that is called reaping. It uses no CPU and no memory, only a PID.

You cannot kill it - it is already dead, so kill -9 does nothing. The bug is in the parent, which is not reaping. Fix or stop the parent, and its zombies are handed to PID 1, which always reaps, so they vanish at once.

A few zombies are cosmetic; thousands mean a parent that will eventually use up every PID, and then nothing new can start. To find the parent: ps -eo pid,ppid,stat,comm, look for Z, and follow the PPID.

Also asked: What is an orphan process, and what happens to it? · What is special about PID 1? · Why does sudo kill -9 1 do nothing?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.