OnCallReady

systemd: interview questions

The question you are most likely to get for each topic, a model answer, and what else comes up. From chapter 2 of the course.

A service is down. How do you find out why? Junior

  1. systemctl status svc: the Active line (failed? activating (auto-restart)? running since 3 seconds ago?), the Result:, the Main PID's exit code, and the last ten log lines.
  2. Read the code. 200 and above is systemd's own (203/EXEC, 217/USER, 200/CHDIR): the program never ran, so fix the unit or the box. A low number is the program's own exit code; code=killed, signal=KILL means it was killed.
  3. journalctl -u svc --since "30 min ago" for the story, and -g 'error|fatal' for the program's own complaints - printed lines are logged at info, so -p err misses them.
  4. journalctl -b -p warning and journalctl -k for what else went wrong on the box.
  5. Check what was loaded: systemctl cat svc and the "changed on disk" warning (forgotten daemon-reload).
  6. Fix the cause, then reset-failed if it hit the start limit, start it, and watch journalctl -u svc -f.

Also asked: Write a minimal unit file for a long-running app and explain each line. · What is the difference between enabling a service and starting it? · How do you change a setting of a service that a package installed?

What are the main sections of a systemd service unit, and what goes in each? Junior

The suffix of the file says the unit type: .service, .timer, .target, .socket, .mount. To look at one: systemctl cat ssh prints the file(s) as written, systemctl show ssh what systemd actually loaded. When they disagree, someone edited the file and forgot daemon-reload.

Also asked: What unit types exist besides services? · What is a target, and what is multi-user.target? · What is the difference between systemctl cat and systemctl show?

Learn it: 2.1 Units and sections

How do you change a setting of a service installed by a package, for example add an environment variable? Junior

With a drop-in: sudo systemctl edit cron, and write only the change, under its section:

[Service]
Environment=LAB=1

That creates /etc/systemd/system/cron.service.d/override.conf and reloads systemd when you save; then sudo systemctl restart cron, because a running process keeps the settings it started with.

Why not edit the file itself: the package's unit lives in /usr/lib/systemd/system/ and apt upgrade replaces it, silently dropping your change. /etc/systemd/system/ wins over /run and /usr/lib, and survives upgrades. Check with systemctl cat cron (shows the drop-in after the main file) and the Drop-In: line in systemctl status. systemctl revert cron undoes it. Gotcha: list settings like ExecStart= add up - clear with an empty ExecStart= first.

Also asked: Which directories does systemd search for unit files, and which one wins? · When would you use systemctl edit --full instead of a drop-in? · You changed a service's settings but nothing happened. What do you check?

Learn it: 2.3 Precedence and drop-ins

How would you turn a program you run by hand into a systemd service? Junior

  1. Put it at a fixed absolute path and make it executable (chmod +x; a script needs its shebang).
  2. Make it run in the foreground and print to stdout - no &, no backgrounding itself, or systemd thinks it exited.
  3. Write /etc/systemd/system/app.service: [Service] with User= (not root), ExecStart=/usr/local/bin/app (absolute path; no pipes, ~ or other shell syntax), Restart=on-failure, RestartSec=10; [Install] with WantedBy=multi-user.target.
  4. sudo systemctl daemon-reload, then sudo systemctl enable --now app: enable = at every boot, --now = also start it now.
  5. Check: systemctl status app and journalctl -u app -f.

What you get over running it by hand: it survives logout and reboot, comes back after a crash, and everything it prints lands in the journal with a timestamp.

Also asked: What does Type=simple mean, and when would you use Type=oneshot? · Why is daemon-reload needed after editing a unit? · Why do services print their logs instead of writing their own log files?

Learn it: 2.5 Your first service

A service fails with status=203/EXEC. What does it mean, and how do you fix it? Junior

203 is systemd's own code: it prepared the process but could not exec the program in ExecStart=. Your program never ran a single line, so its own logs are empty.

The journal says why, read bottom-up: Failed at step EXEC spawning /path: Permission denied. The usual causes, in order:

  1. not executable - chmod +x;
  2. wrong path or a typo - No such file or directory;
  3. a missing or broken shebang (a blank line above it, or Windows line endings);
  4. the interpreter is not where the shebang says;
  5. a sandbox setting hides it - a script in /home with ProtectHome=yes.

Before reloading a hand-written unit, systemd-analyze verify catches most of these. Its neighbours: 217/USER (the User= does not exist) and 200/CHDIR (no such WorkingDirectory=).

Also asked: What does an exit status of 200 or higher mean in systemctl status? · What does systemd-analyze verify check? · Why does a ~ in ExecStart= give a "bad-setting" error?

Learn it: 2.7 status=203/EXEC

What is the difference between a service being enabled and being active? Junior

Enabled is about boot; active is about right now. They are independent.

Enabled means systemctl enable created a symlink (in multi-user.target.wants/, from the [Install] section), so the unit starts at the next boot. It starts nothing now. Active means it is running at this moment.

All four combinations happen: enabled and running (normal), enabled but dead (crashed or stopped), disabled but running (started by hand - it will not come back after a reboot), disabled and dead. ssh on Ubuntu is "disabled and running", because ssh.socket starts it on the first connection.

systemctl status shows both: Loaded: ...; enabled; and Active: active (running). In a script: systemctl is-enabled svc and systemctl is-active svc; enable --now does both at once.

Also asked: Walk me through the lines of systemctl status. · Why is checking $? after systemctl start not proof that the service started? · What is the difference between inactive (dead) and failed?

Learn it: 2.9 Reading systemctl status, and the exit-code table

A service has Restart=on-failure. Why does kill PID leave it stopped while kill -9 PID brings it back? Junior

Because Restart= looks at how the process ended.

kill PID sends SIGTERM (15), "please finish up and exit". systemd counts SIGTERM, SIGINT, SIGHUP and SIGPIPE as clean terminations, so the unit goes to inactive (dead) and the journal says "Deactivated successfully" - no restart. kill -9 sends SIGKILL, which is unclean, so on-failure applies: activating (auto-restart), then a new PID after RestartSec=.

The values: no (default), on-success, on-failure (non-zero exit, unclean signal, timeout, watchdog), on-abnormal, on-abort, always - and none of them restarts after an explicit systemctl stop. A program that exits 143 after SIGTERM needs SuccessExitStatus=143, or every clean stop counts as a failure. Check with systemctl show -p Restart and -p MainPID --value.

Also asked: What Restart= values are there, and which would you choose for a web app? · What does RestartSec= do, and why is the default too fast? · Why would a service need SuccessExitStatus=143?

Learn it: 2.10 Restart policy and clean signals

A service shows "start request repeated too quickly". What does it mean, and what do you do? Junior

It hit its start limit: more than StartLimitBurst starts within StartLimitIntervalSec (default 5 in 10 seconds). systemd gave up and marked it failed (Result: start-limit-hit); it will not start again by itself - that is the brake on a crash loop.

What I do: systemctl status svc for the last exit code, then journalctl -u svc for the program's own error, usually the same line over and over (a missing config file, a port already in use). Fix that, then sudo systemctl reset-failed svc and start it again. reset-failed fixes nothing by itself; it only lets you try.

Two gotchas: the StartLimit* settings belong in [Unit] - in [Service] they are ignored with only a warning from systemd-analyze verify. And with RestartSec=10 the default 10-second window can never hold five starts, so the limit never trips.

Also asked: What is a crash loop, and how does systemd stop one? · How do you list every failed unit on a box? · What does StartLimitIntervalSec=0 do, and when would a service want it?

Learn it: 2.12 The start limit, and where it lives

What is the difference between After= and Requires= (or Wants=)? Junior

They are independent: one is about time, the other about existence.

So you almost always want a pair: After=postgresql.service with Requires= or Wants=. The classic case is the network: After=network.target does not mean an address is assigned. Use After=network-online.target and Wants=network-online.target. Check with systemctl list-dependencies and systemctl show -p After -p Wants.

Also asked: A service fails at boot but works when you start it by hand. What would you suspect? · What is the difference between network.target and network-online.target? · What do BindsTo=, PartOf= and Conflicts= do?

Learn it: 2.14 Ordering is not dependency

How would you schedule a script to run every day at 02:00 with systemd? Junior

Two units with the same base name.

cleanup.service says what to run: [Service] with Type=oneshot and ExecStart=/usr/local/bin/cleanup.sh, and no [Install] section.

cleanup.timer says when:

[Timer]
OnCalendar=*-*-* 02:00:00
Persistent=true
RandomizedDelaySec=10m

[Install]
WantedBy=timers.target

Then sudo systemctl daemon-reload && sudo systemctl enable --now cleanup.timer - you enable the timer, never the service. Persistent=true runs a missed run at the next boot if the box was off at 02:00; the random delay stops a whole fleet firing at the same second. Check the expression with systemd-analyze calendar '*-*-* 02:00:00', the schedule with systemctl list-timers, and the output with journalctl -u cleanup.service.

Also asked: Why use a systemd timer instead of a cron line? · What is the difference between OnCalendar= and OnUnitActiveSec=? · A nightly job did not run. Where do you look?

Learn it: 2.16 Timers, the cron replacement

What is the difference between disabling a service and masking it? Junior

systemctl disable only removes the boot symlinks that enable created. The unit still starts by hand, or when something else wants it - so "I disabled it" does not mean "it is off".

systemctl mask points the unit at /dev/null with a symlink in /etc, so it cannot start at all: not by hand, not as a dependency, not at boot. Any attempt fails with "Unit cron.service is masked." unmask undoes it. I mask when I need a guarantee, like maintenance where an automatic start would do harm.

And the boot side: enable = start at the next boot, start = run now, enable --now = both. systemctl is-enabled prints one word: enabled, disabled, masked, or static (no [Install] section; other units pull it in - not "off").

Also asked: What does static mean in systemctl is-enabled? · What are targets, and how do you change the one a server boots into? · Why is systemctl isolate rescue.target dangerous over SSH?

Learn it: 2.18 enable, disable, mask, static, targets

How should a systemd service get a database password, and why not with Environment=? Junior

Not with Environment=: environment variables are not a secret store. systemctl show -p Environment svc prints them to any user, no privileges needed, and root can read them in /proc/PID/environ.

EnvironmentFile= is a little better, because you can lock the file with chmod 600 - but the value still ends up in /proc/PID/environ.

The systemd-native way is LoadCredential=db-password:/etc/creds/db-password. systemd, as root, reads the file and gives the service a private copy at $CREDENTIALS_DIRECTORY/db-password: kept in memory, readable only by the service's user, not in the environment, and gone when the unit stops. The program reads its password from that file. The principle: a file only the service can read, never an environment variable.

Also asked: What are ExecStartPre=, ExecReload= and ExecStop= used for? · What does the - prefix do in EnvironmentFile=-/etc/default/app? · You changed Environment= and ran daemon-reload, but the program still sees the old value. Why?

Learn it: 2.21 Exec directives, environment and secrets

What happens when you run systemctl stop on a service? Junior

  1. systemd runs ExecStop=, if there is one.
  2. It sends SIGTERM to the main process and, with the default KillMode=control-group, to every process in the service's cgroup.
  3. It waits up to TimeoutStopSec=, 90 seconds by default - the grace period for the program to finish cleanly.
  4. Anything still alive gets SIGKILL.

SIGTERM can be caught: a program with a signal handler closes connections, saves its work and exits. SIGKILL cannot be caught; the process is gone with whatever it was doing - half-written files, cut-off requests. A program that ignores SIGTERM takes the full 90 seconds, the journal says State 'stop-sigterm' timed out. Killing., and the unit ends failed with result timeout. A shorter TimeoutStopSec= in a drop-in limits the wait, but it is a mitigation: the real fix is handling SIGTERM.

Also asked: What is the difference between SIGTERM and SIGKILL? · Why does a reboot sometimes hang on "A stop job is running"? · What does KillMode= control?

Learn it: 2.24 Shutdown and stop timeouts

How do you limit how much memory and CPU a service may use? Junior

systemd puts every service in its own cgroup, and the kernel enforces limits on the whole group. In the unit (or a drop-in):

At runtime, without editing files: sudo systemctl set-property demo MemoryMax=100M CPUQuota=20% - saved under /etc/systemd/system.control/, so it outlives a reboot. To check, read the kernel's own view: cat /sys/fs/cgroup/system.slice/demo.service/memory.max (max = no limit) and memory.current; systemd-cgtop shows usage live.

Also asked: What does ProtectSystem=strict do, and what do you add so the service can still write its data? · Why does a script in /home fail with 203/EXEC once ProtectHome=yes is set? · What does systemd-analyze security tell you?

Learn it: 2.26 Hardening and cgroup limits

What is a template unit like [email protected], and when would you use one? Junior

A unit file with @ before the suffix is a template: one file systemd stamps out as many times as you ask. You never start [email protected] itself; you start instances - sudo systemctl start worker@alpha worker@beta - and inside the file the part after the @ is available as %i:

[Service]
ExecStart=/usr/local/bin/worker --queue %i

Use it when you need several copies of the same thing that differ only by a name: four workers for four queues, without four near-identical files drifting apart. Ubuntu itself does it: [email protected] is the console login prompt, [email protected] your per-user systemd. systemctl list-units 'worker@*' shows the running instances; %n is the full unit name, handy in OnFailure=notify-oncall@%n.service.

Also asked: How can a service tell someone when it fails, without extra monitoring software? · What does a watchdog (WatchdogSec=) catch that Restart= alone does not? · What is %i, and what are specifiers?

Learn it: 2.28 Templates and failure handling

A server rebooted unexpectedly overnight. How do you find out why? Junior

The answer is in the log of the boot before this one; the current boot's log starts after the fact.

  1. journalctl --list-boots - is the previous boot there at all? Index 0 is this boot, -1 the one before.
  2. journalctl -b -1 -e - jump to the end of the previous boot: the last lines before it went down.
  3. journalctl -b -1 -p err - errors and worse (0-3) from that boot, and journalctl -b -1 -k for the kernel's messages.

The trap: if /var/log/journal does not exist, the journal is volatile - kept in /run/log/journal in RAM - and every reboot wipes it, so -b -1 finds nothing and the reason is gone. That is why I make it persistent with sudo mkdir -p /var/log/journal and sudo journalctl --flush (or Storage=persistent), and keep its size in check with --vacuum-time=7d or SystemMaxUse=.

Also asked: Which journalctl options do you use most, and for what? · What does journalctl -p err include, and what does it miss? · How do you stop the journal from filling the disk?

Learn it: 2.30 journald: priorities, boots and persistence

A service failed at around 14:05. How do you find the cause in the logs? Junior

I ask the journal precise questions instead of scrolling, always with a time window:

  1. systemctl status X - state, exit code, last lines.
  2. journalctl -u X --since "13:55" --until "14:10" --no-pager - the story around the failure.
  3. journalctl -u X --since "13:55" -g 'error|exception|fatal' - the program's own complaint. -p err would miss it: whatever a service prints is logged at info.
  4. journalctl -b -p warning --since "13:55" - what else on the box was going wrong; often the service failed because something else did.
  5. journalctl -k --since "13:55" - the kernel: memory kills, disk errors.

To see what a noisy service says most: journalctl -u X -o cat | sort | uniq -c | sort -rn | head. For a ticket, -o short-iso gives sortable, unambiguous timestamps.

Also asked: What does "structured logging" mean in the journal, and why does it matter? · What is the difference between _SYSTEMD_UNIT= and UNIT= matches? · How do you show two services' logs interleaved in time order?

Learn it: 2.32 journalctl as a query language

A VM takes two minutes to boot. How do you find out why? Junior

systemd-analyze time first: it splits the boot into the kernel part and the userspace part (the units). On a server it is almost always userspace.

Then systemd-analyze critical-chain, read bottom-up: the chain of units each waiting for the one below, with @ when each finished and + how long it took. That chain is what actually delayed the boot. The most common answer is systemd-networkd-wait-online.service, the service that makes network-online.target true, waiting for a network.

systemd-analyze blame lists units by start time, slowest first - but it misleads on its own: units start in parallel, so a slow unit nothing waits for costs nothing. Use blame to find candidates, critical-chain to find the one that matters.

Also asked: What is the difference between daemon-reload, daemon-reexec and systemctl restart? · Why can systemd-analyze blame mislead you? · What does systemctl revert undo?

Learn it: 2.33 systemd-analyze and the boot

Practise these answers with flashcards and labs Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.