OnCallReady

Lesson 2.26 · systemd · 14 min read

Hardening and cgroup limits

In plain words

Imagine letting a new cleaner into your house. You would not give them the keys to every room, the safe and the car. You give them one key, lock the study, and tell them the rubbish goes in one bin. And you give them a set time and a set budget, so one cleaner cannot use the whole day or the whole cupboard of supplies.

Hardening directives are the locked rooms: ProtectSystem=strict makes the files read-only for that service, ProtectHome=yes hides /home, PrivateTmp=yes gives it a private /tmp, NoNewPrivileges=yes forbids picking up extra keys. MemoryMax, CPUQuota and TasksMax are the budget, enforced by the service's cgroup. systemd-analyze security demo scores how many rooms are still open.

Why this matters

Two bad days this lesson prevents. One: a bug in your web app lets an attacker run commands - as your service's user, with access to everything that user can reach. Two: a memory leak (a program that keeps using more memory and never gives it back) in one service slowly eats all the RAM and takes every other service on the box down with it. systemd can put each service in a smaller room, with a lock on the door and a limit on how much it can use.

What you need to know already: 2.3 (drop-ins), 2.7 (203/EXEC), 2.24 (the cgroup: a service and all its processes, grouped by the kernel).

Sandboxing is a few lines, not a project

Hardening (or sandboxing) means taking away everything a service does not need, so a broken or attacked service can do less damage. Each directive below removes one thing:

[Service]
NoNewPrivileges=yes         # can never gain more power than it started with
PrivateTmp=yes              # its own private /tmp, invisible to others
ProtectSystem=strict        # the entire filesystem read-only...
ReadWritePaths=/var/lib/app # ...except here
ProtectHome=yes             # /home, /root, /run/user hidden completely
PrivateDevices=yes          # no direct access to disks and other hardware
ProtectKernelTunables=yes   # cannot change kernel settings under /proc/sys, /sys
ProtectKernelModules=yes    # cannot load kernel add-ons (modules)
RestrictSUIDSGID=yes        # cannot create files that run with extra power
LockPersonality=yes         # cannot switch into odd compatibility modes

NoNewPrivileges matters because some programs (like sudo) are marked so that they run with their owner's power, often root's; with this setting such a program started by the service gets no extra power. /tmp is the shared scratch folder every user and program can write into.

ProtectSystem= has three levels: yes (/usr and /boot read-only), full (plus /etc), strict (the whole filesystem, except /dev, /proc and /sys).

The gotcha that will bite you: ProtectHome=yes makes /home inaccessible to the service (ProtectHome=tmpfs makes it look empty instead, and read-only keeps it readable). A unit whose ExecStart points at a script in /home/learner/... then fails with 203/EXEC - Permission denied - and no amount of chmod helps, because that process is not allowed into /home at all. The fix is to put the script where a service's programs should live: /usr/local/bin/ (the folder for programs installed by hand, not by apt).

DynamicUser=yes goes further: systemd invents a temporary user account (a fresh UID, user ID number) for as long as the unit runs, with PrivateTmp, ProtectHome=read-only and ProtectSystem=strict switched on automatically. Nothing to create, nothing to clean up. Only usable if the service does not need to own files that stay on disk.

Scoring it

systemd-analyze security demo

Lists every sandbox setting, whether it is set, and ends with an exposure score from 0 (locked down) to 10 (wide open) on the last line (Overall exposure level for demo.service: ...). A unit with none of these settings scores 9.x. Under ~7 is respectable for a normal service; below 2 usually means DynamicUser plus a full set of Protect* lines.

Do not chase the number blindly - read what each line costs you.

Resource limits

The cgroup from 2.24 is also how the kernel limits a service: how much memory (RAM) it may use, how much CPU time, how many processes. You set the limits in the unit:

MemoryMax=512M      hard memory limit. Go over it and the kernel kills the
                    service's process (SIGKILL, exit code 137). The part of
                    the kernel that does that is the OOM killer (OOM = out of
                    memory). Chapter 5 is all about it.
MemoryHigh=400M     soft limit: above this the kernel slows the service down
                    and takes memory back, but does not kill. Set it below
                    MemoryMax as an early warning.
CPUQuota=20%        20% of ONE CPU core (a core = one processor that can run
                    one thing at a time). 200% means two full cores.
TasksMax=64         max processes/threads (a thread is a lighter process
                    inside a process). A runaway program that keeps starting
                    copies of itself hits this, not the whole box.
IOWeight=50         priority for disk reads/writes (IO) relative to other
                    services, 1-10000, default 100.

At runtime, without editing any file:

sudo systemctl set-property demo MemoryMax=100M CPUQuota=20%

which is saved in /etc/systemd/system.control/demo.service.d/ - a real directory you can look at, and the reason a limit you do not remember setting can outlive a reboot.

You can also run a one-off command inside its own temporary unit, with limits: sudo systemd-run --scope -p MemoryMax=200M some-command (--scope = run it right here in your terminal, -p = set a property). Handy for testing what a limit does.

Seeing the cgroups

systemd-cgls                 the cgroup tree, unit by unit, with their processes
systemd-cgtop                like a live dashboard, per cgroup: CPU, memory, IO
cat /sys/fs/cgroup/system.slice/demo.service/memory.max
cat /sys/fs/cgroup/system.slice/demo.service/memory.current

/sys/fs/cgroup is a folder the kernel fills with one directory per cgroup; the files inside are the kernel's own view of the limits (memory.max, in bytes: 100M = 104857600) and current usage (memory.current). system.slice is the group systemd puts all system services in. systemd writes those files for you; the kernel enforces them.

Later (Ch 10): containers are limited through exactly these cgroup files, so what you learn here carries straight over.

What you can now do

Why it helps

Hardening is expected more and more in reviews and audits: a service with ProtectSystem=strict and NoNewPrivileges that gets broken into still cannot rewrite /etc or gain root. Adding a handful of lines to a unit is one of the cheapest security wins there is, and systemd-analyze security gives you a measurable before and after.

Resource limits pay off during incidents: a memory leak in one service gets that service killed at its MemoryMax instead of the kernel picking some other victim - maybe the SSH server you need to fix things. CPUQuota stops one runaway job from making the whole box sluggish. The cgroup files you read here (memory.max, memory.current) are the kernel's real numbers, not systemd's opinion.

Commands in this lesson

systemd-analyze cat

FAQ

Why does my service fail with 203/EXEC after adding ProtectHome=yes?

ProtectHome=yes makes /home, /root and /run/user invisible to the service. If ExecStart points at a script under /home/learner, the service cannot reach it and launching fails with Permission denied, which no chmod can fix. The right fix is to install the program where services belong, /usr/local/bin or /opt, rather than removing the protection.

What is the difference between MemoryMax and MemoryHigh?

MemoryMax is the hard limit: if the service cannot be squeezed back under it, the kernel kills a process in its cgroup (Result: oom-kill, "out of memory"). MemoryHigh is a brake: above it, the kernel slows the service down and takes memory back where it can, but does not kill. Setting MemoryHigh a little below MemoryMax gives an early, visible slowdown before a kill.

What does CPUQuota=20% mean on a machine with two cores?

20% of one core's time. The service may use at most 0.2 seconds of processor time per second in total, across all cores; CPUQuota=200% allows two full cores. When it hits the quota it is made to wait (throttled), which shows up as the service getting slower, not as an error - so a too-tight quota looks like a latency problem, not a crash.

Where does systemctl set-property store its changes?

It applies the change immediately and saves it as a drop-in under /etc/systemd/system.control/unit.d/, which has even higher priority than /etc/systemd/system. That is why a limit can survive reboots even though nobody edited a unit file. --runtime makes the change temporary instead (kept under /run, gone at reboot). systemctl cat unit shows these drop-ins too.

Is a low systemd-analyze security score always the goal?

No. The score counts protections that are switched on, but some do not fit a given service: a backup agent must read /home, and a service that keeps files between runs needs a fixed user rather than DynamicUser. Use the list as a checklist: switch on what the service does not need, then test it properly. A locked-down unit that breaks in production helps nobody.

In an interview Junior

How do you limit how much memory and CPU a service may use?

systemd puts every service in its own cgroup, and the kernel enforces limits on the whole group. In the unit (or a drop-in):

At runtime, without editing files: sudo systemctl set-property demo MemoryMax=100M CPUQuota=20% - saved under /etc/systemd/system.control/, so it outlives a reboot. To check, read the kernel's own view: cat /sys/fs/cgroup/system.slice/demo.service/memory.max (max = no limit) and memory.current; systemd-cgtop shows usage live.

Also asked: What does ProtectSystem=strict do, and what do you add so the service can still write its data? · Why does a script in /home fail with 203/EXEC once ProtectHome=yes is set? · What does systemd-analyze security tell you?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.