Chapter 2 systemd
Units, drop-ins, restarts, dependencies, timers, hardening, resource limits, the journal - and an agent that will not stay up.
In plain words
Imagine a theatre with a stage manager who holds a clipboard for every performer. Each sheet says who the performer is, who must be on stage before them, what to do if they faint, and how long to wait before walking them off when the show ends. The stage manager starts everyone in the right order, helps up anyone who collapses, and writes everything that happens in one big logbook.
The stage manager is systemd, the first program the kernel starts (PID 1). The clipboard sheets are unit files like demo.service and heartbeat.timer. systemctl is how you talk to the stage manager, and journalctl is how you read the logbook. This chapter is about writing those sheets well, and reading the logbook when a performer keeps fainting.
Why it matters on call
Almost every program that runs on a Linux server - the SSH server you log in through, the web server, the scheduled cleanup job, your team's own app - is started, watched and restarted by systemd. When something is "down", the first two commands anyone types are systemctl status <name> and journalctl -u <name>. Being fluent in them is what lets you answer "is it running, since when, and what did it say before it died?" in thirty seconds.
It also connects straight to Chapter 0. Restart policies and failure alerts are how you keep an SLO; the journal is where the timeline of a postmortem comes from; replacing a hand-run fix with a timer is removing toil. And "write a unit file for this app" or "your service keeps restarting, debug it" are among the most common hands-on Linux interview tasks.
Lessons
- Units and sections
- Precedence and drop-ins
- Your first service
- status=203/EXEC
- Reading systemctl status, and the exit-code table
- Restart policy and clean signals
- The start limit, and where it lives
- Ordering is not dependency
- Timers, the cron replacement
- enable, disable, mask, static, targets
- Exec directives, environment and secrets
- Shutdown and stop timeouts
- Hardening and cgroup limits
- Templates and failure handling
- journald: priorities, boots and persistence
- journalctl as a query language
- systemd-analyze and the boot
22 hands-on labs (missions, incidents and drills) run in the terminal: Open this chapter in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.
Questions people ask
Is systemd replacing cron, syslog and old init scripts?
On Ubuntu, largely yes. Old start scripts in /etc/init.d still work through a compatibility layer, but packages ship unit files now. Timers can replace cron (the older job scheduler), though cron is still installed and used. journald collects all logs and also passes them on to rsyslog, which still writes classic text files like /var/log/syslog. You will meet both worlds, so learn to read each.
What is the difference between systemctl and service?
service is the older command from before systemd. On a systemd machine, service nginx restart simply calls systemctl restart nginx. It still works and shows up in old instructions, but it hides systemd features like the detailed status, enable, mask, show and drop-ins. Use systemctl directly; it is what all current documentation assumes.
Where do unit files live and which one wins?
Three directories, in order of priority: /etc/systemd/system (yours, the admin's), /run/systemd/system (temporary, gone at reboot) and /usr/lib/systemd/system (installed by packages). The first directory that has a file with that name provides the whole unit, and drop-ins in name.service.d/*.conf are added on top. systemctl cat name shows exactly which files are in use, so you never have to guess.
Do I need to memorise all the directives?
No. Learn the twenty or so that appear in almost every unit (Description, After, Wants, Type, ExecStart, User, Restart, RestartSec, WantedBy...) and know where to look for the rest: man systemd.service, man systemd.exec and man systemd.unit. systemd-analyze verify tells you when you have mistyped one or put it in the wrong section.
Why does this chapter matter for interviews?
Because it is one of the most common hands-on Linux topics. Expect "write a unit file for this app", "your service keeps restarting, how do you debug it", "what is the difference between Wants and Requires", "how would you make logs survive a reboot". Answering with the exact commands (systemctl status, journalctl -u -b -1, systemd-analyze verify) is what separates practice from reading.