systemd: interview questions
The question you are most likely to get for each topic, a model answer, and what else comes up. From chapter 2 of the course.
A service is down. How do you find out why? Junior
systemctl status svc: the Active line (failed?activating (auto-restart)? running since 3 seconds ago?), theResult:, the Main PID's exit code, and the last ten log lines.- Read the code. 200 and above is systemd's own (203/EXEC, 217/USER, 200/CHDIR): the program never ran, so fix the unit or the box. A low number is the program's own exit code;
code=killed, signal=KILLmeans it was killed. journalctl -u svc --since "30 min ago"for the story, and-g 'error|fatal'for the program's own complaints - printed lines are logged at info, so-p errmisses them.journalctl -b -p warningandjournalctl -kfor what else went wrong on the box.- Check what was loaded:
systemctl cat svcand the "changed on disk" warning (forgottendaemon-reload). - Fix the cause, then
reset-failedif it hit the start limit, start it, and watchjournalctl -u svc -f.
Also asked: Write a minimal unit file for a long-running app and explain each line. · What is the difference between enabling a service and starting it? · How do you change a setting of a service that a package installed?
What are the main sections of a systemd service unit, and what goes in each? Junior
- [Unit]: what it is and how it relates to other units -
Description=, ordering and dependencies (After=,Wants=), and the start limit (StartLimitIntervalSec=,StartLimitBurst=). - [Service]: how to run the program -
Type=,ExecStart=(the exact command, absolute path),User=,WorkingDirectory=,Restart=,RestartSec=. - [Install]: what
systemctl enableshould do -WantedBy=multi-user.targetmeans "start me at boot". Nothing in it happens until you enable the unit.
The suffix of the file says the unit type: .service, .timer, .target, .socket, .mount. To look at one: systemctl cat ssh prints the file(s) as written, systemctl show ssh what systemd actually loaded. When they disagree, someone edited the file and forgot daemon-reload.
Also asked: What unit types exist besides services? · What is a target, and what is multi-user.target? · What is the difference between systemctl cat and systemctl show?
Learn it: 2.1 Units and sections
How do you change a setting of a service installed by a package, for example add an environment variable? Junior
With a drop-in: sudo systemctl edit cron, and write only the change, under its section:
[Service]
Environment=LAB=1
That creates /etc/systemd/system/cron.service.d/override.conf and reloads systemd when you save; then sudo systemctl restart cron, because a running process keeps the settings it started with.
Why not edit the file itself: the package's unit lives in /usr/lib/systemd/system/ and apt upgrade replaces it, silently dropping your change. /etc/systemd/system/ wins over /run and /usr/lib, and survives upgrades. Check with systemctl cat cron (shows the drop-in after the main file) and the Drop-In: line in systemctl status. systemctl revert cron undoes it. Gotcha: list settings like ExecStart= add up - clear with an empty ExecStart= first.
Also asked: Which directories does systemd search for unit files, and which one wins? · When would you use systemctl edit --full instead of a drop-in? · You changed a service's settings but nothing happened. What do you check?
Learn it: 2.3 Precedence and drop-ins
How would you turn a program you run by hand into a systemd service? Junior
- Put it at a fixed absolute path and make it executable (
chmod +x; a script needs its shebang). - Make it run in the foreground and print to stdout - no
&, no backgrounding itself, or systemd thinks it exited. - Write
/etc/systemd/system/app.service:[Service]withUser=(not root),ExecStart=/usr/local/bin/app(absolute path; no pipes,~or other shell syntax),Restart=on-failure,RestartSec=10;[Install]withWantedBy=multi-user.target. sudo systemctl daemon-reload, thensudo systemctl enable --now app: enable = at every boot,--now= also start it now.- Check:
systemctl status appandjournalctl -u app -f.
What you get over running it by hand: it survives logout and reboot, comes back after a crash, and everything it prints lands in the journal with a timestamp.
Also asked: What does Type=simple mean, and when would you use Type=oneshot? · Why is daemon-reload needed after editing a unit? · Why do services print their logs instead of writing their own log files?
Learn it: 2.5 Your first service
A service fails with status=203/EXEC. What does it mean, and how do you fix it? Junior
203 is systemd's own code: it prepared the process but could not exec the program in ExecStart=. Your program never ran a single line, so its own logs are empty.
The journal says why, read bottom-up: Failed at step EXEC spawning /path: Permission denied. The usual causes, in order:
- not executable -
chmod +x; - wrong path or a typo - No such file or directory;
- a missing or broken shebang (a blank line above it, or Windows line endings);
- the interpreter is not where the shebang says;
- a sandbox setting hides it - a script in
/homewithProtectHome=yes.
Before reloading a hand-written unit, systemd-analyze verify catches most of these. Its neighbours: 217/USER (the User= does not exist) and 200/CHDIR (no such WorkingDirectory=).
Also asked: What does an exit status of 200 or higher mean in systemctl status? · What does systemd-analyze verify check? · Why does a ~ in ExecStart= give a "bad-setting" error?
Learn it: 2.7 status=203/EXEC
What is the difference between a service being enabled and being active? Junior
Enabled is about boot; active is about right now. They are independent.
Enabled means systemctl enable created a symlink (in multi-user.target.wants/, from the [Install] section), so the unit starts at the next boot. It starts nothing now. Active means it is running at this moment.
All four combinations happen: enabled and running (normal), enabled but dead (crashed or stopped), disabled but running (started by hand - it will not come back after a reboot), disabled and dead. ssh on Ubuntu is "disabled and running", because ssh.socket starts it on the first connection.
systemctl status shows both: Loaded: ...; enabled; and Active: active (running). In a script: systemctl is-enabled svc and systemctl is-active svc; enable --now does both at once.
Also asked: Walk me through the lines of systemctl status. · Why is checking $? after systemctl start not proof that the service started? · What is the difference between inactive (dead) and failed?
Learn it: 2.9 Reading systemctl status, and the exit-code table
A service has Restart=on-failure. Why does kill PID leave it stopped while kill -9 PID brings it back? Junior
Because Restart= looks at how the process ended.
kill PID sends SIGTERM (15), "please finish up and exit". systemd counts SIGTERM, SIGINT, SIGHUP and SIGPIPE as clean terminations, so the unit goes to inactive (dead) and the journal says "Deactivated successfully" - no restart. kill -9 sends SIGKILL, which is unclean, so on-failure applies: activating (auto-restart), then a new PID after RestartSec=.
The values: no (default), on-success, on-failure (non-zero exit, unclean signal, timeout, watchdog), on-abnormal, on-abort, always - and none of them restarts after an explicit systemctl stop. A program that exits 143 after SIGTERM needs SuccessExitStatus=143, or every clean stop counts as a failure. Check with systemctl show -p Restart and -p MainPID --value.
Also asked: What Restart= values are there, and which would you choose for a web app? · What does RestartSec= do, and why is the default too fast? · Why would a service need SuccessExitStatus=143?
Learn it: 2.10 Restart policy and clean signals
A service shows "start request repeated too quickly". What does it mean, and what do you do? Junior
It hit its start limit: more than StartLimitBurst starts within StartLimitIntervalSec (default 5 in 10 seconds). systemd gave up and marked it failed (Result: start-limit-hit); it will not start again by itself - that is the brake on a crash loop.
What I do: systemctl status svc for the last exit code, then journalctl -u svc for the program's own error, usually the same line over and over (a missing config file, a port already in use). Fix that, then sudo systemctl reset-failed svc and start it again. reset-failed fixes nothing by itself; it only lets you try.
Two gotchas: the StartLimit* settings belong in [Unit] - in [Service] they are ignored with only a warning from systemd-analyze verify. And with RestartSec=10 the default 10-second window can never hold five starts, so the limit never trips.
Also asked: What is a crash loop, and how does systemd stop one? · How do you list every failed unit on a box? · What does StartLimitIntervalSec=0 do, and when would a service want it?
Learn it: 2.12 The start limit, and where it lives
What is the difference between After= and Requires= (or Wants=)? Junior
They are independent: one is about time, the other about existence.
After=foo: ordering only. If foo is also starting, wait for it. It does not pull foo in, and does not care whether foo succeeded.Requires=foo: a hard dependency. Starting me starts foo; stopping or restarting foo stops me too. On its own it says nothing about order, so both start at the same instant - a race.Wants=foo: a soft dependency. Try to start foo, carry on regardless. The most common and safest.
So you almost always want a pair: After=postgresql.service with Requires= or Wants=. The classic case is the network: After=network.target does not mean an address is assigned. Use After=network-online.target and Wants=network-online.target. Check with systemctl list-dependencies and systemctl show -p After -p Wants.
Also asked: A service fails at boot but works when you start it by hand. What would you suspect? · What is the difference between network.target and network-online.target? · What do BindsTo=, PartOf= and Conflicts= do?
Learn it: 2.14 Ordering is not dependency
How would you schedule a script to run every day at 02:00 with systemd? Junior
Two units with the same base name.
cleanup.service says what to run: [Service] with Type=oneshot and ExecStart=/usr/local/bin/cleanup.sh, and no [Install] section.
cleanup.timer says when:
[Timer]
OnCalendar=*-*-* 02:00:00
Persistent=true
RandomizedDelaySec=10m
[Install]
WantedBy=timers.target
Then sudo systemctl daemon-reload && sudo systemctl enable --now cleanup.timer - you enable the timer, never the service. Persistent=true runs a missed run at the next boot if the box was off at 02:00; the random delay stops a whole fleet firing at the same second. Check the expression with systemd-analyze calendar '*-*-* 02:00:00', the schedule with systemctl list-timers, and the output with journalctl -u cleanup.service.
Also asked: Why use a systemd timer instead of a cron line? · What is the difference between OnCalendar= and OnUnitActiveSec=? · A nightly job did not run. Where do you look?
Learn it: 2.16 Timers, the cron replacement
What is the difference between disabling a service and masking it? Junior
systemctl disable only removes the boot symlinks that enable created. The unit still starts by hand, or when something else wants it - so "I disabled it" does not mean "it is off".
systemctl mask points the unit at /dev/null with a symlink in /etc, so it cannot start at all: not by hand, not as a dependency, not at boot. Any attempt fails with "Unit cron.service is masked." unmask undoes it. I mask when I need a guarantee, like maintenance where an automatic start would do harm.
And the boot side: enable = start at the next boot, start = run now, enable --now = both. systemctl is-enabled prints one word: enabled, disabled, masked, or static (no [Install] section; other units pull it in - not "off").
Also asked: What does static mean in systemctl is-enabled? · What are targets, and how do you change the one a server boots into? · Why is systemctl isolate rescue.target dangerous over SSH?
How should a systemd service get a database password, and why not with Environment=? Junior
Not with Environment=: environment variables are not a secret store. systemctl show -p Environment svc prints them to any user, no privileges needed, and root can read them in /proc/PID/environ.
EnvironmentFile= is a little better, because you can lock the file with chmod 600 - but the value still ends up in /proc/PID/environ.
The systemd-native way is LoadCredential=db-password:/etc/creds/db-password. systemd, as root, reads the file and gives the service a private copy at $CREDENTIALS_DIRECTORY/db-password: kept in memory, readable only by the service's user, not in the environment, and gone when the unit stops. The program reads its password from that file. The principle: a file only the service can read, never an environment variable.
Also asked: What are ExecStartPre=, ExecReload= and ExecStop= used for? · What does the - prefix do in EnvironmentFile=-/etc/default/app? · You changed Environment= and ran daemon-reload, but the program still sees the old value. Why?
What happens when you run systemctl stop on a service? Junior
- systemd runs
ExecStop=, if there is one. - It sends SIGTERM to the main process and, with the default
KillMode=control-group, to every process in the service's cgroup. - It waits up to
TimeoutStopSec=, 90 seconds by default - the grace period for the program to finish cleanly. - Anything still alive gets SIGKILL.
SIGTERM can be caught: a program with a signal handler closes connections, saves its work and exits. SIGKILL cannot be caught; the process is gone with whatever it was doing - half-written files, cut-off requests. A program that ignores SIGTERM takes the full 90 seconds, the journal says State 'stop-sigterm' timed out. Killing., and the unit ends failed with result timeout. A shorter TimeoutStopSec= in a drop-in limits the wait, but it is a mitigation: the real fix is handling SIGTERM.
Also asked: What is the difference between SIGTERM and SIGKILL? · Why does a reboot sometimes hang on "A stop job is running"? · What does KillMode= control?
Learn it: 2.24 Shutdown and stop timeouts
How do you limit how much memory and CPU a service may use? Junior
systemd puts every service in its own cgroup, and the kernel enforces limits on the whole group. In the unit (or a drop-in):
MemoryMax=512M- a hard limit: go over it and the kernel's OOM killer kills the process (SIGKILL, exit code 137).MemoryHigh=400M- a soft limit: the service is slowed down and memory taken back, but not killed.CPUQuota=20%- 20% of one core; 200% is two cores.TasksMax=64- max processes and threads.
At runtime, without editing files: sudo systemctl set-property demo MemoryMax=100M CPUQuota=20% - saved under /etc/systemd/system.control/, so it outlives a reboot. To check, read the kernel's own view: cat /sys/fs/cgroup/system.slice/demo.service/memory.max (max = no limit) and memory.current; systemd-cgtop shows usage live.
Also asked: What does ProtectSystem=strict do, and what do you add so the service can still write its data? · Why does a script in /home fail with 203/EXEC once ProtectHome=yes is set? · What does systemd-analyze security tell you?
Learn it: 2.26 Hardening and cgroup limits
What is a template unit like [email protected], and when would you use one? Junior
A unit file with @ before the suffix is a template: one file systemd stamps out as many times as you ask. You never start [email protected] itself; you start instances - sudo systemctl start worker@alpha worker@beta - and inside the file the part after the @ is available as %i:
[Service]
ExecStart=/usr/local/bin/worker --queue %i
Use it when you need several copies of the same thing that differ only by a name: four workers for four queues, without four near-identical files drifting apart. Ubuntu itself does it: [email protected] is the console login prompt, [email protected] your per-user systemd. systemctl list-units 'worker@*' shows the running instances; %n is the full unit name, handy in OnFailure=notify-oncall@%n.service.
Also asked: How can a service tell someone when it fails, without extra monitoring software? · What does a watchdog (WatchdogSec=) catch that Restart= alone does not? · What is %i, and what are specifiers?
Learn it: 2.28 Templates and failure handling
A server rebooted unexpectedly overnight. How do you find out why? Junior
The answer is in the log of the boot before this one; the current boot's log starts after the fact.
journalctl --list-boots- is the previous boot there at all? Index 0 is this boot, -1 the one before.journalctl -b -1 -e- jump to the end of the previous boot: the last lines before it went down.journalctl -b -1 -p err- errors and worse (0-3) from that boot, andjournalctl -b -1 -kfor the kernel's messages.
The trap: if /var/log/journal does not exist, the journal is volatile - kept in /run/log/journal in RAM - and every reboot wipes it, so -b -1 finds nothing and the reason is gone. That is why I make it persistent with sudo mkdir -p /var/log/journal and sudo journalctl --flush (or Storage=persistent), and keep its size in check with --vacuum-time=7d or SystemMaxUse=.
Also asked: Which journalctl options do you use most, and for what? · What does journalctl -p err include, and what does it miss? · How do you stop the journal from filling the disk?
A service failed at around 14:05. How do you find the cause in the logs? Junior
I ask the journal precise questions instead of scrolling, always with a time window:
systemctl status X- state, exit code, last lines.journalctl -u X --since "13:55" --until "14:10" --no-pager- the story around the failure.journalctl -u X --since "13:55" -g 'error|exception|fatal'- the program's own complaint.-p errwould miss it: whatever a service prints is logged at info.journalctl -b -p warning --since "13:55"- what else on the box was going wrong; often the service failed because something else did.journalctl -k --since "13:55"- the kernel: memory kills, disk errors.
To see what a noisy service says most: journalctl -u X -o cat | sort | uniq -c | sort -rn | head. For a ticket, -o short-iso gives sortable, unambiguous timestamps.
Also asked: What does "structured logging" mean in the journal, and why does it matter? · What is the difference between _SYSTEMD_UNIT= and UNIT= matches? · How do you show two services' logs interleaved in time order?
Learn it: 2.32 journalctl as a query language
A VM takes two minutes to boot. How do you find out why? Junior
systemd-analyze time first: it splits the boot into the kernel part and the userspace part (the units). On a server it is almost always userspace.
Then systemd-analyze critical-chain, read bottom-up: the chain of units each waiting for the one below, with @ when each finished and + how long it took. That chain is what actually delayed the boot. The most common answer is systemd-networkd-wait-online.service, the service that makes network-online.target true, waiting for a network.
systemd-analyze blame lists units by start time, slowest first - but it misleads on its own: units start in parallel, so a slow unit nothing waits for costs nothing. Use blame to find candidates, critical-chain to find the one that matters.
Also asked: What is the difference between daemon-reload, daemon-reexec and systemctl restart? · Why can systemd-analyze blame mislead you? · What does systemctl revert undo?
Learn it: 2.33 systemd-analyze and the boot
Practise these answers with flashcards and labs Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.