OnCallReady

Lesson 2.10 · systemd · 12 min read

Restart policy and clean signals

In plain words

Imagine a nurse with instructions for a patient who sometimes faints. "Help them up if they faint" is not the same as "drag them back if they leave". If the patient says goodbye politely and walks out, the nurse should let them go.

Restart=on-failure is that instruction. A polite goodbye is a clean stop: exit 0, or one of the signals systemd treats as clean - SIGTERM, SIGINT, SIGHUP and SIGPIPE. kill <PID> sends SIGTERM, so the service stays down. kill -9 is fainting: SIGKILL is not clean, so systemd waits RestartSec and starts it again with a new PID. systemctl stop always wins; nobody drags you back after that.

Why this matters

"The service crashed and systemd brought it back" and "I stopped it and it stayed dead" are both things you want. systemd decides between them by looking at how the process ended. Get this wrong and either a crash leaves your app down, or a deliberate stop keeps coming back to life.

What you need to know already: 2.5 (demo.service, Restart=on-failure), 2.9 (reading systemctl status), exit codes (1.7).

Signals, just enough

A signal is a short message the kernel delivers to a process, usually to ask it to stop. Each has a name and a number. Two matter here:

The kill command sends a signal to a PID: kill 1234 sends SIGTERM, kill -9 1234 sends SIGKILL. (Despite the name, kill just sends signals. Chapter 3 covers all of them.)

Restart= is about why it stopped

no            (default) never restart
on-success    only on a clean exit
on-failure    non-zero exit, an unclean signal, a timeout, or a watchdog
on-abnormal   signal / timeout / watchdog, but NOT a non-zero exit code
on-abort      only an uncaught signal
always        whatever happened - a clean exit, a crash, a signal. The one
              thing it does NOT override is an explicit "systemctl stop".

A clean exit means the program finished by itself with exit code 0 (success). RestartSec= is the pause before trying again (default 100ms - far too fast for anything that depends on another machine being back).

The part that surprises people: which signals count as clean

systemd treats SIGTERM, SIGINT, SIGHUP and SIGPIPE as clean terminations. So with Restart=on-failure:

kill <PID>        -> SIGTERM -> clean   -> inactive (dead). NO restart.
kill -9 <PID>     -> SIGKILL -> unclean -> restart, after RestartSec.

That is why "I killed it and systemd brought it back" and "I killed it and it stayed dead" are both true, depending on which signal you used. Watch the Active line:

Active: inactive (dead)                          <- after SIGTERM
Active: activating (auto-restart) (Result: signal) <- after SIGKILL, waiting
Active: active (running)                         <- after RestartSec elapsed

and the journal says Deactivated successfully for the clean case.

SuccessExitStatus

Some programs exit non-zero on purpose. Many catch SIGTERM, clean up, and then exit with code 143 - by convention 128 + the signal number (15), meaning "I ended because of SIGTERM". systemd would call that a failure. So you tell it otherwise:

SuccessExitStatus=143

Now a 143 counts as success and Restart=on-failure leaves it alone.

Scripting around it

systemctl show -p MainPID --value demo    # just the number
systemctl is-active demo                  # one word, exit 0 if active
systemctl is-failed demo

is-active is the one to use in a script - it prints one word and sets its exit code (0 = active), where status output is formatted for humans and changes between versions.

What you can now do

Why it helps

Restart behaviour decides whether a crash at 3am is a blip nobody notices or a page. You will be asked both "why did this service not come back after it was killed?" and "why does it keep coming back when I kill it?", and the answer is always the combination of Restart= and which signal was used.

It also explains two real annoyances: programs that exit with a non-zero code on every normal stop and get recorded as failed (fixed with SuccessExitStatus=), and a very short RestartSec that turns a database blip into a storm of reconnect attempts from every service at once. In SLO terms (Chapter 0), a good restart policy shortens the time to recover.

Commands in this lesson

kill systemctl

FAQ

Should I use Restart=always or Restart=on-failure?

on-failure for most services: restart after crashes, unclean signals, timeouts and watchdog failures, but respect a clean exit. always also restarts after exit 0 and clean signals - useful for agents that must never stay down, or programs that exit 0 on purpose after a while. Both leave the service stopped after an explicit systemctl stop.

Why does SIGTERM count as clean?

SIGTERM is the polite "please stop" request - the very signal systemd itself sends on systemctl stop. A process that ends because of it did what it was asked, so systemd does not treat it as a crash. SIGINT (what Ctrl+C sends), SIGHUP and SIGPIPE are also clean by default. SIGKILL and crash signals are unclean. SuccessExitStatus= adds more codes or signals to the clean list.

How do I see how many times a service restarted?

systemctl show svc -p NRestarts --value gives the count since the unit was last started by hand. The journal has one line per attempt, "Scheduled restart job, restart counter is at N", so journalctl -u svc | grep -c 'Scheduled restart' counts them over a time window (grep -c counts matching lines). The Active: ... since time tells you when the current run began.

Does RestartSec grow if the service keeps failing?

Not by default: RestartSec is a fixed pause, so systemd retries at the same pace until the start limit (2.12) trips. Recent systemd versions can make the pause grow with RestartSteps= and RestartMaxDelaySec=, so a service that keeps failing is retried less and less often instead of hammering whatever it depends on.

Why does kill without -9 not bring my service back?

Because plain kill sends SIGTERM, which systemd treats as a clean stop, and Restart=on-failure does not restart after a clean stop. The service goes to inactive (dead) with "Deactivated successfully" in the journal. kill -9 sends SIGKILL, which is unclean, so it restarts. If you actually want a restart, use systemctl restart.

In an interview Junior

A service has Restart=on-failure. Why does kill PID leave it stopped while kill -9 PID brings it back?

Because Restart= looks at how the process ended.

kill PID sends SIGTERM (15), "please finish up and exit". systemd counts SIGTERM, SIGINT, SIGHUP and SIGPIPE as clean terminations, so the unit goes to inactive (dead) and the journal says "Deactivated successfully" - no restart. kill -9 sends SIGKILL, which is unclean, so on-failure applies: activating (auto-restart), then a new PID after RestartSec=.

The values: no (default), on-success, on-failure (non-zero exit, unclean signal, timeout, watchdog), on-abnormal, on-abort, always - and none of them restarts after an explicit systemctl stop. A program that exits 143 after SIGTERM needs SuccessExitStatus=143, or every clean stop counts as a failure. Check with systemctl show -p Restart and -p MainPID --value.

Also asked: What Restart= values are there, and which would you choose for a web app? · What does RestartSec= do, and why is the default too fast? · Why would a service need SuccessExitStatus=143?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.