Why this matters
"The service crashed and systemd brought it back" and "I stopped it and it stayed dead" are both things you want. systemd decides between them by looking at how the process ended. Get this wrong and either a crash leaves your app down, or a deliberate stop keeps coming back to life.
What you need to know already: 2.5 (demo.service, Restart=on-failure), 2.9 (reading systemctl status), exit codes (1.7).
Signals, just enough
A signal is a short message the kernel delivers to a process, usually to ask it to stop. Each has a name and a number. Two matter here:
- SIGTERM (15, "terminate") - "please finish up and exit". The polite one; the program can clean up first.
- SIGKILL (9, "kill") - "you are gone, now". The kernel removes the process immediately; it gets no say.
The kill command sends a signal to a PID: kill 1234 sends SIGTERM, kill -9 1234 sends SIGKILL. (Despite the name, kill just sends signals. Chapter 3 covers all of them.)
Restart= is about why it stopped
no (default) never restart
on-success only on a clean exit
on-failure non-zero exit, an unclean signal, a timeout, or a watchdog
on-abnormal signal / timeout / watchdog, but NOT a non-zero exit code
on-abort only an uncaught signal
always whatever happened - a clean exit, a crash, a signal. The one
thing it does NOT override is an explicit "systemctl stop".
A clean exit means the program finished by itself with exit code 0 (success). RestartSec= is the pause before trying again (default 100ms - far too fast for anything that depends on another machine being back).
The part that surprises people: which signals count as clean
systemd treats SIGTERM, SIGINT, SIGHUP and SIGPIPE as clean terminations. So with Restart=on-failure:
kill <PID> -> SIGTERM -> clean -> inactive (dead). NO restart.
kill -9 <PID> -> SIGKILL -> unclean -> restart, after RestartSec.
That is why "I killed it and systemd brought it back" and "I killed it and it stayed dead" are both true, depending on which signal you used. Watch the Active line:
Active: inactive (dead) <- after SIGTERM
Active: activating (auto-restart) (Result: signal) <- after SIGKILL, waiting
Active: active (running) <- after RestartSec elapsed
and the journal says Deactivated successfully for the clean case.
SuccessExitStatus
Some programs exit non-zero on purpose. Many catch SIGTERM, clean up, and then exit with code 143 - by convention 128 + the signal number (15), meaning "I ended because of SIGTERM". systemd would call that a failure. So you tell it otherwise:
SuccessExitStatus=143
Now a 143 counts as success and Restart=on-failure leaves it alone.
Scripting around it
systemctl show -p MainPID --value demo # just the number
systemctl is-active demo # one word, exit 0 if active
systemctl is-failed demo
is-active is the one to use in a script - it prints one word and sets its exit code (0 = active), where status output is formatted for humans and changes between versions.
What you can now do
- predict whether systemd restarts a service after SIGTERM vs SIGKILL
- pick a
Restart=value, and useSuccessExitStatus=for odd exit codes - get a service's PID or state in one word for a script