OnCallReady

Lesson 2.9 · systemd · 26 min read

Reading systemctl status, and the exit-code table

In plain words

Think of the display panel on a washing machine. One glance tells you which programme it is running, whether it is on, how long it has been going, whether it paused, and the last few beeps it made. If it says "cycle started 3 seconds ago" but you loaded it an hour ago, something keeps restarting it.

systemctl status is that panel for a service. The dot colour and the Active: line say what state it is in and since when. Loaded: says which unit file won and whether it starts at boot. Main PID: says how the last run ended. Exit codes of 200 and up are systemd's own ("could not even load the machine"); lower numbers are the program's own.

The problem

A service is misbehaving and someone asks you "what is wrong with it?". You type systemctl status and get fifteen dense lines. Every answer you need is on that screen - whether it runs, since when, which file it came from, how it last died - but only if you can read it. This lesson reads it line by line.

What you need to know already: units and unit files (2.1), drop-ins (2.3), your first service and Type= (2.5), what 203/EXEC means (2.7), exit codes (Driving the shell, Chapter 1).

A few words first

One screen, many answers

● ssh.service - OpenBSD Secure Shell server
     Loaded: loaded (/usr/lib/systemd/system/ssh.service; disabled; preset: enabled)
    Drop-In: /etc/systemd/system/ssh.service.d
             └─override.conf
     Active: active (running) since Tue 2026-09-22 17:43:01 UTC; 2h 17min ago
 Invocation: 4d2c31d466a14744445444819a51e27a
TriggeredBy: ● ssh.socket
       Docs: man:sshd(8)
             man:sshd_config(5)
    Process: 17380 ExecStartPre=/usr/sbin/sshd -t (code=exited, status=0/SUCCESS)
   Main PID: 700 (sshd)
      Tasks: 1 (limit: 4583)
     Memory: 8.1M (peak: 8.1M)
        CPU: 4119ms
     CGroup: /system.slice/ssh.service
             └─700 "sshd: /usr/sbin/sshd -D [listener] 0 of 10-100 startups"

Sep 22 17:43:01 oncall-lab sshd[700]: Server listening on 0.0.0.0 port 22.

Top to bottom:

The dot. ● green = active (running or done). ● white = in the middle of starting or stopping. ○ = inactive. × red = failed. On a long list-units screen you can spot trouble by colour alone.

Loaded: three facts in one line.

bad-setting or not-found on this line means systemd never even tried to run the unit.

Drop-In: every drop-in file that was merged in (2.3). If "my change is not taking effect", either it is missing here or systemd has not reloaded.

Active: the state, the sub-state in brackets (the finer detail: running, exited, dead...), and since when. The "since" is the most underrated field on the screen: active (running) since ... 3s ago on a service that should have been up for weeks means it keeps dying and restarting.

Invocation: a fresh random ID every time the unit starts. Every journal line from that run carries it, so on a real box journalctl _SYSTEMD_INVOCATION_ID=4d2c... shows exactly one run.

TriggeredBy: what starts this unit on demand. It only appears when a .socket or .timer unit (2.1) starts it. Since Ubuntu 22.10, ssh works this way: ssh.socket holds port 22, and ssh.service is disabled and only starts on the first connection - your login. That is why it says disabled and is running. (It also means a new Port in sshd's config needs sudo systemctl daemon-reload && sudo systemctl restart ssh.socket.)

Docs: where the manual pages are. man sshd would open the first one.

Process: helper commands that already finished - ExecStartPre= and friends - with their exit status. A failed pre-start check shows up here, not on the Main PID line.

Main PID: the one process systemd is watching over. For a failed unit this line keeps the last one and how it ended: Main PID: 17398 (code=exited, status=3/NOTIMPLEMENTED).

Tasks / Memory / CPU: totals for the whole service's cgroup, not just one process. Tasks = how many processes and threads; (limit: 4583) is the maximum allowed (the TasksMax= setting). CPU is total processor time used so far.

CGroup: the tree of processes inside the service. Everything listed here is stopped when you stop the unit.

The log lines at the bottom are the last ten journal entries for the unit - the same as journalctl -u ssh -n 10, without typing it.

Active states you must recognise

active (running)              a daemon, up
active (exited)               a oneshot with RemainAfterExit=yes that finished
active (waiting)              a timer, waiting for its next run
activating (start-pre)        ExecStartPre= is still running
activating (auto-restart)     it died; systemd waits RestartSec, then starts it again
deactivating (stop-sigterm)   systemd sent SIGTERM and is waiting for it to exit
inactive (dead)               stopped, or never started, or exited cleanly
failed (Result: ...)          it stopped, and the way it stopped counts as a failure

Restart=, RestartSec= and the waiting after SIGTERM come in 2.10 and 2.24; for now, just recognise the words.

The Result= values

When a unit fails, Result: says how:

exit-code          the main process exited with a non-zero exit code
signal             killed by a signal systemd does not treat as a clean stop (SIGKILL...)
core-dump          it crashed, and the kernel saved a copy of its memory for debugging
timeout            starting or stopping took longer than allowed
oom-kill           the kernel killed it for using more memory than it was allowed
start-limit-hit    it was started too many times in a short window; systemd gave up (2.12)
resources          systemd could not even prepare the process (a missing EnvironmentFile...)
watchdog           the program promised regular "I am alive" pings and they stopped (2.28)
exec-condition     an ExecCondition= check said "do not run"

systemctl show -p Result --value unit prints just the word, for scripts.

The exit-code table

An exit code (from Driving the shell) is the number a program returns when it ends: 0 = success, anything else = some kind of failure.

Codes 200 and above are systemd's own. They mean the failure happened while systemd was preparing the process, before your program ran its first line. Your application's logs are empty, because your application never ran.

200/CHDIR          WorkingDirectory= does not exist or cannot be entered
203/EXEC           ExecStart= could not be run: missing file, not executable, bad #!
                   line, missing interpreter, or hidden by a sandbox setting (2.26)
209/STDOUT         the StandardOutput= target could not be opened
214/SETSCHEDULER   CPUSchedulingPolicy= failed
217/USER           the User= or Group= does not exist
226/NAMESPACE      a sandbox setting (2.26) could not be set up - often a
                   ReadWritePaths= directory that does not exist
243/CREDENTIALS    a LoadCredential= source file is missing (2.21)

Below 200 it is your program's own exit code. systemd prints a name next to it from LSB (Linux Standard Base, an old standard that named codes 1-7), and that is where confusing output like this comes from:

Main process exited, code=exited, status=3/NOTIMPLEMENTED

The program did not claim anything was "not implemented". It exited with 3, and LSB happens to call 3 NOTIMPLEMENTED. The names: 1 FAILURE, 2 INVALIDARGUMENT, 3 NOTIMPLEMENTED, 4 NOPERMISSION, 5 NOTINSTALLED, 6 NOTCONFIGURED, 7 NOTRUNNING. Read the number, then read the program's own last log lines above it.

A death by signal is reported as a signal, not an exit code. The journal line says Main process exited, code=killed, status=9/KILL (signal number 9 is SIGKILL); systemctl status says Main PID: 17406 (code=killed, signal=KILL). The shell has its own habit of reporting "killed by signal N" as exit code 128+N (so 137 for signal 9). systemd does not do that: in its output, 9 is the signal number.

Three failures side by side

These are example units, not ones on this box. Read the Active and Main PID lines of each.

× u217.service - bad user
     Active: failed (Result: exit-code) since Tue 2026-09-22 20:00:03 UTC; 100ms ago
   Main PID: 17387 (code=exited, status=217/USER)
... u217.service: Failed to determine user credentials: No such process
... u217.service: Failed at step USER spawning /usr/local/bin/ok.sh: No such process

A User= that does not exist: 217, and "Failed at step USER".

× uenv.service - bad envfile
     Active: failed (Result: resources) since Tue 2026-09-22 20:00:03 UTC; 100ms ago
... uenv.service: Failed to load environment files: No such file or directory
... uenv.service: Failed to spawn 'start' task: No such file or directory

EnvironmentFile= (a file of NAME=value settings handed to the program, 2.21) is missing: Result: resources, and no Main PID at all.

○ urel.service - relative
     Loaded: bad-setting (Reason: Unit urel.service has a bad unit file setting.)
     Active: inactive (dead)
... urel.service: Service has no ExecStart=, ExecStop=, or SuccessAction=. Refusing.

Note the third: Loaded: bad-setting and inactive (dead), not failed. The unit never started. The first journal line about it (Executable path is not absolute, ignoring) was logged when systemd loaded the file, at daemon-reload time, not when you tried to start it - so it is further up in the journal than you might look.

"Job for X failed" versus silence

# example transcript: uenv and u217 are the broken units above, not units on this box
sudo systemctl start uenv
Job for uenv.service failed because the control process exited with error code.
See "systemctl status uenv.service" and "journalctl -xeu uenv.service" for details.
sudo systemctl start u217
echo $?
0

(echo $? prints the exit code of the previous command.) The second start "succeeded" even though the unit failed. Why: with Type=simple (the default, 2.5) the start is done the moment systemd has created the process. If running the program then fails, systemctl start has already returned 0. So a deploy script that runs systemctl start and checks $? has proved nothing. Two fixes:

Scriptable status: systemctl show

systemctl status is for humans. For scripts, systemctl show prints the unit's properties (its settings and live values) as Name=value lines. -p picks which properties; --value drops the Name= part.

# ucrash: a unit that exits 3 on every start (example, not on this box)
systemctl show ucrash -p ActiveState,SubState,Result,ExecMainStatus,NRestarts
ActiveState=failed
SubState=failed
Result=start-limit-hit
ExecMainStatus=3
NRestarts=5
$ systemctl show -p MainPID --value demo
18342

ExecMainStatus is the main process's exit code; NRestarts counts automatic restarts. Never search through systemctl status text in a script - it is laid out for people and changes between versions. show -p properties are the stable interface.

"changed on disk"

Warning: The unit file, source configuration file or drop-ins of uenv.service changed on disk. Run 'systemctl daemon-reload' to reload units.

systemd runs the copy of the unit it loaded into memory, not the file you just saved. Every systemctl command about the unit prints this until you run daemon-reload. If you see it, nothing you edited has taken effect yet - not even the fix.

What you can now do

Why it helps

This is the screen you read in the first 30 seconds of every service incident, and reading all of it is a skill. "since 3s ago" on a service that should have run for weeks means a restart loop. Result: start-limit-hit means systemd gave up. bad-setting means the unit never ran. The Drop-In line reveals overrides someone forgot about.

The scripting half matters just as much. A deploy script that runs systemctl start and checks $? has proved nothing for a Type=simple unit; systemctl is-active and systemctl show -p are the reliable way to ask. That is exactly the kind of fix you would suggest when reviewing a deploy script, or write down as an action item after an incident.

Commands in this lesson

systemctl

FAQ

Why did systemctl start return 0 when the service failed?

With Type=simple, starting is finished as soon as systemd has created the process. If launching the program then fails, or it crashes a second later, systemctl start has already returned success (0). Use Type=exec so launch failures fail the command, and always check afterwards with systemctl is-active or, better, by asking the app itself whether it works.

What does status=3/NOTIMPLEMENTED mean? My app has no such idea.

Nothing specific. The program exited with code 3, and systemd printed the name an old standard (LSB) gave that number. LSB named codes 1 to 7 (1 FAILURE, 2 INVALIDARGUMENT, 3 NOTIMPLEMENTED...), but almost no program follows it. Read the number, then read the program's own last log lines above it, which explain why it exited.

What is the difference between inactive (dead) and failed?

Both mean not running. inactive (dead) means it stopped in a way systemd considers normal: never started, stopped by you, or exited cleanly (exit 0, or a polite SIGTERM). failed means the way it ended counts as a failure: a non-zero exit, a kill, a timeout, the start limit. Failed units show in systemctl --failed, and systemctl reset-failed clears that state.

Why should I not search the text of systemctl status in scripts?

It is written for people: cut to the width of your terminal, coloured, mixed with log lines, and its layout changes between systemd versions (the Invocation: line is new, for example). systemctl is-active, is-enabled and is-failed give clean exit codes, and systemctl show unit -p ActiveState,SubState,Result --value gives stable Name=value output meant for programs.

What does code=killed, status=9/KILL mean compared with exit code 137?

The same event described two ways. systemd reports that the process was killed by signal number 9 (SIGKILL). A shell reports the same death as exit code 137, using its habit of showing "killed by signal N" as 128 + N. So in systemd's output, 9 is a signal number, not an exit code. The Result: word tells you who killed it, for example timeout for systemd's own stop timeout.

In an interview Junior

What is the difference between a service being enabled and being active?

Enabled is about boot; active is about right now. They are independent.

Enabled means systemctl enable created a symlink (in multi-user.target.wants/, from the [Install] section), so the unit starts at the next boot. It starts nothing now. Active means it is running at this moment.

All four combinations happen: enabled and running (normal), enabled but dead (crashed or stopped), disabled but running (started by hand - it will not come back after a reboot), disabled and dead. ssh on Ubuntu is "disabled and running", because ssh.socket starts it on the first connection.

systemctl status shows both: Loaded: ...; enabled; and Active: active (running). In a script: systemctl is-enabled svc and systemctl is-active svc; enable --now does both at once.

Also asked: Walk me through the lines of systemctl status. · Why is checking $? after systemctl start not proof that the service started? · What is the difference between inactive (dead) and failed?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.