Why this matters
Two bad days this lesson prevents. One: a bug in your web app lets an attacker run commands - as your service's user, with access to everything that user can reach. Two: a memory leak (a program that keeps using more memory and never gives it back) in one service slowly eats all the RAM and takes every other service on the box down with it. systemd can put each service in a smaller room, with a lock on the door and a limit on how much it can use.
What you need to know already: 2.3 (drop-ins), 2.7 (203/EXEC), 2.24 (the cgroup: a service and all its processes, grouped by the kernel).
Sandboxing is a few lines, not a project
Hardening (or sandboxing) means taking away everything a service does not need, so a broken or attacked service can do less damage. Each directive below removes one thing:
[Service]
NoNewPrivileges=yes # can never gain more power than it started with
PrivateTmp=yes # its own private /tmp, invisible to others
ProtectSystem=strict # the entire filesystem read-only...
ReadWritePaths=/var/lib/app # ...except here
ProtectHome=yes # /home, /root, /run/user hidden completely
PrivateDevices=yes # no direct access to disks and other hardware
ProtectKernelTunables=yes # cannot change kernel settings under /proc/sys, /sys
ProtectKernelModules=yes # cannot load kernel add-ons (modules)
RestrictSUIDSGID=yes # cannot create files that run with extra power
LockPersonality=yes # cannot switch into odd compatibility modes
NoNewPrivileges matters because some programs (like sudo) are marked so that they run with their owner's power, often root's; with this setting such a program started by the service gets no extra power. /tmp is the shared scratch folder every user and program can write into.
ProtectSystem= has three levels: yes (/usr and /boot read-only), full (plus /etc), strict (the whole filesystem, except /dev, /proc and /sys).
The gotcha that will bite you: ProtectHome=yes makes /home inaccessible to the service (ProtectHome=tmpfs makes it look empty instead, and read-only keeps it readable). A unit whose ExecStart points at a script in /home/learner/... then fails with 203/EXEC - Permission denied - and no amount of chmod helps, because that process is not allowed into /home at all. The fix is to put the script where a service's programs should live: /usr/local/bin/ (the folder for programs installed by hand, not by apt).
DynamicUser=yes goes further: systemd invents a temporary user account (a fresh UID, user ID number) for as long as the unit runs, with PrivateTmp, ProtectHome=read-only and ProtectSystem=strict switched on automatically. Nothing to create, nothing to clean up. Only usable if the service does not need to own files that stay on disk.
Scoring it
systemd-analyze security demo
Lists every sandbox setting, whether it is set, and ends with an exposure score from 0 (locked down) to 10 (wide open) on the last line (Overall exposure level for demo.service: ...). A unit with none of these settings scores 9.x. Under ~7 is respectable for a normal service; below 2 usually means DynamicUser plus a full set of Protect* lines.
Do not chase the number blindly - read what each line costs you.
Resource limits
The cgroup from 2.24 is also how the kernel limits a service: how much memory (RAM) it may use, how much CPU time, how many processes. You set the limits in the unit:
MemoryMax=512M hard memory limit. Go over it and the kernel kills the
service's process (SIGKILL, exit code 137). The part of
the kernel that does that is the OOM killer (OOM = out of
memory). Chapter 5 is all about it.
MemoryHigh=400M soft limit: above this the kernel slows the service down
and takes memory back, but does not kill. Set it below
MemoryMax as an early warning.
CPUQuota=20% 20% of ONE CPU core (a core = one processor that can run
one thing at a time). 200% means two full cores.
TasksMax=64 max processes/threads (a thread is a lighter process
inside a process). A runaway program that keeps starting
copies of itself hits this, not the whole box.
IOWeight=50 priority for disk reads/writes (IO) relative to other
services, 1-10000, default 100.
At runtime, without editing any file:
sudo systemctl set-property demo MemoryMax=100M CPUQuota=20%
which is saved in /etc/systemd/system.control/demo.service.d/ - a real directory you can look at, and the reason a limit you do not remember setting can outlive a reboot.
You can also run a one-off command inside its own temporary unit, with limits: sudo systemd-run --scope -p MemoryMax=200M some-command (--scope = run it right here in your terminal, -p = set a property). Handy for testing what a limit does.
Seeing the cgroups
systemd-cgls the cgroup tree, unit by unit, with their processes
systemd-cgtop like a live dashboard, per cgroup: CPU, memory, IO
cat /sys/fs/cgroup/system.slice/demo.service/memory.max
cat /sys/fs/cgroup/system.slice/demo.service/memory.current
/sys/fs/cgroup is a folder the kernel fills with one directory per cgroup; the files inside are the kernel's own view of the limits (memory.max, in bytes: 100M = 104857600) and current usage (memory.current). system.slice is the group systemd puts all system services in. systemd writes those files for you; the kernel enforces them.
Later (Ch 10): containers are limited through exactly these cgroup files, so what you learn here carries straight over.
What you can now do
- sandbox a service and measure it with
systemd-analyze security - recognise the ProtectHome 203/EXEC trap
- cap a service's memory and CPU, and read the limit back from /sys/fs/cgroup