OnCallReady

Lesson 2.24 · systemd · 9 min read

Shutdown and stop timeouts

In plain words

Imagine the end of the school day. First the bell rings politely: "time to pack up". Most children finish their sentence, pack their bags and leave. The teacher waits a few minutes. Anyone still sitting there after that gets walked out of the room, finished or not.

systemctl stop works the same way. It runs ExecStop= if there is one, sends SIGTERM (the polite bell) to every process in the service's cgroup, waits TimeoutStopSec= (90 seconds by default), then sends SIGKILL, which cannot be refused. A program that ignores the bell makes everyone wait the full time and ends up failed (Result: timeout).

Why this matters

You type sudo reboot and the machine sits there for a minute and a half: "A stop job is running for ...". Or you stop a service and it takes forever, then the program loses the work it was doing. Both come from the same few rules about how systemd stops things.

What you need to know already: 2.10 (SIGTERM vs SIGKILL, Restart=), 2.12 (list-units --failed, reset-failed), 2.21 (ExecStop=).

What systemctl stop actually does

A program often starts helper processes of its own, called child processes (the program is their parent). systemd keeps a service and all its children together in one group, its cgroup (control group: a kernel feature that bundles processes so they can be counted, limited and signalled together - more in 2.26).

  1. Run ExecStop= if there is one.
  2. Send SIGTERM (or whatever KillSignal= says) to the main process - and, unless KillMode=process, to every process in the unit's cgroup.
  3. Wait up to TimeoutStopSec= (default 90 seconds).
  4. If anything is still alive, send SIGKILL.

That wait is the grace period: time for the program to finish cleanly. A program can choose to ignore SIGTERM; one that does takes the full 90 seconds to stop, and the journal shows:

stubborn.service: State 'stop-sigterm' timed out. Killing.
stubborn.service: Killing process 1234 (stubborn) with signal SIGKILL.
stubborn.service: Main process exited, code=killed, status=9/KILL
stubborn.service: Failed with result 'timeout'.

Note it ends up failed, not inactive. A unit that times out on stop is a failure, and it will show in systemctl list-units --failed.

That 90 seconds is also why a reboot sometimes hangs for a minute and a half with "A stop job is running for ...". The fix is never to reboot harder; it is to find out why the process will not exit.

KillMode

control-group  (default) signal every process in the cgroup
mixed          SIGTERM to the main process only, SIGKILL to all of them
process        signal only the main process - children are left running
none           signal nothing (do not)

control-group is what you want almost always: it is the guarantee that a service which starts child processes does not leave them running after it stops (leftover processes whose parent is gone are called orphans). process is for daemons that manage their own children and would be upset by systemd killing them - cron uses it, so a long-running cron job survives a cron restart.

Why the program must handle SIGTERM

To handle SIGTERM, a program contains a signal handler: a bit of code that runs when the signal arrives, finishes the current work and exits. Without one, every stop - every deploy (installing a new version) and every reboot - ends in SIGKILL, and whatever the program was in the middle of is lost. For a web server that means requests cut off halfway: users see errors (HTTP 502, "bad gateway") for a moment after each deploy, and nobody connects the two.

Lowering the timeout does not fix a service that ignores SIGTERM - it just makes you kill it sooner. It is a mitigation (it limits the damage), not a repair.

Later (Ch 15): Kubernetes stops containers with exactly this SIGTERM, wait, SIGKILL sequence; its grace period is called terminationGracePeriodSeconds.

What you can now do

Why it helps

Slow or messy shutdowns cause real incidents: deploys that drop requests that were half-way through and show a burst of errors, reboots that hang for 90 seconds on "A stop job is running", programs killed in the middle of writing a file. Understanding the SIGTERM, wait, SIGKILL sequence lets you find the real cause - the program ignores SIGTERM, or a child process of it does - instead of just shortening the timeout.

The same three-step pattern (ask politely, wait a fixed time, force it) is how almost every system that runs programs stops them, so learning it here pays off again and again.

Later (Ch 15): Kubernetes stops the programs it runs with exactly this sequence; its grace period is TimeoutStopSec.

Commands in this lesson

systemctl

FAQ

Why does a stop timeout end in failed rather than inactive?

Because the service did not stop the way it was asked to. systemd had to escalate to SIGKILL, so the unit's result is timeout and its state becomes failed. It shows up in systemctl --failed, can fire an alert, and must be cleared with reset-failed. That is useful: a service that needs SIGKILL to stop has a bug worth tracking.

Does ExecStop replace the SIGTERM?

No, it runs first. systemd runs the ExecStop= commands, then sends KillSignal= (SIGTERM by default) to whatever processes are still in the cgroup, waits TimeoutStopSec, then sends SIGKILL. If ExecStop already made the service exit cleanly, there is nothing left to signal. That order makes ExecStop the place for "finish current work, then exit" commands.

Should I just lower TimeoutStopSec for a service that stops slowly?

Only as a stopgap. If the service ignores SIGTERM, a lower timeout just kills it sooner, still losing half-finished work, and still marks it failed. Find out why it ignores SIGTERM: a wrapper script that does not pass the signal on, a program with no handler for it, or a process stuck waiting. Some services really need longer, like a database saving its data, and those need a bigger timeout, not a smaller one.

What does KillMode=process do and why would cron use it?

With KillMode=process, only the main process gets the stop signals; other processes in the cgroup keep running. cron uses it so that restarting cron itself does not kill a long job it had started. The price is that those child processes can outlive the service. The default, control-group, signals every process in the cgroup, so nothing is left behind.

Why does my reboot hang with "A stop job is running"?

During shutdown, systemd stops every unit and waits up to each unit's stop timeout. A service that ignores SIGTERM holds the shutdown for its full timeout, often 90 seconds or more. The message names the unit. After the reboot, read its last lines from the previous boot with journalctl -b -1 -u unit (2.30), and fix its signal handling or its timeout.

In an interview Junior

What happens when you run systemctl stop on a service?

  1. systemd runs ExecStop=, if there is one.
  2. It sends SIGTERM to the main process and, with the default KillMode=control-group, to every process in the service's cgroup.
  3. It waits up to TimeoutStopSec=, 90 seconds by default - the grace period for the program to finish cleanly.
  4. Anything still alive gets SIGKILL.

SIGTERM can be caught: a program with a signal handler closes connections, saves its work and exits. SIGKILL cannot be caught; the process is gone with whatever it was doing - half-written files, cut-off requests. A program that ignores SIGTERM takes the full 90 seconds, the journal says State 'stop-sigterm' timed out. Killing., and the unit ends failed with result timeout. A shorter TimeoutStopSec= in a drop-in limits the wait, but it is a mitigation: the real fix is handling SIGTERM.

Also asked: What is the difference between SIGTERM and SIGKILL? · Why does a reboot sometimes hang on "A stop job is running"? · What does KillMode= control?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.