Why this matters
All containers on a host share one kernel. If an attacker gets code running in a container, what they can do next depends on the flags you ran it with. A few flags turn "they own one app" into "they own nothing much" - and a few others turn it into "they own the host". This lesson is the checklist, with a reason for each item.
What you need to know already: root vs normal users and setuid binaries (4.9); systemd hardening like
NoNewPrivileges=(2.26);/proc/PID/status(3.14); running images as non-root withUSERand the docker group = root (10.40, 10.1).
Root in a container is not all of root
Linux splits root's powers into about 40 separate permissions called capabilities: CAP_CHOWN (change file owners), CAP_NET_RAW (raw network packets, used by ping), CAP_SYS_ADMIN (mounting and much else)... A process can hold some and not others. Root in a default container keeps only 14:
$ docker run --rm alpine:3.20 grep Cap /proc/1/status
CapInh: 0000000000000000
CapPrm: 00000000a80425fb
CapEff: 00000000a80425fb
CapBnd: 00000000a80425fb
CapAmb: 0000000000000000
$ capsh --decode=00000000a80425fb
0x00000000a80425fb=cap_chown,cap_dac_override,cap_fowner,cap_fsetid,cap_kill,cap_setgid,cap_setuid,cap_setpcap,cap_net_bind_service,cap_net_raw,cap_sys_chroot,cap_mknod,cap_audit_write,cap_setfcap
The kernel stores the sets as hex bitmasks, one bit per capability. The one that counts is CapEff (effective: what the process can use right now); CapBnd (bounding) is the most it could ever get. capsh --decode=HEX turns a mask into names (capsh is in the libcap2-bin package, on the host).
Missing on purpose: SYS_ADMIN (mount, most of "root"), NET_ADMIN (change interfaces and firewall rules), SYS_PTRACE (attach to other processes, like strace), SYS_MODULE (load kernel modules), SYS_TIME... So a root container cannot, say, load a kernel module.
A process running as a non-root user has an empty effective set (CapEff: 0000000000000000) whatever the bounding set says - another reason USER matters.
Drop everything, add back what you need
$ docker run --rm --cap-drop ALL alpine:3.20 grep CapEff /proc/1/status
CapEff: 0000000000000000
$ docker run --rm --cap-drop ALL alpine:3.20 chown 1000 /etc/hostname
chown: /etc/hostname: Operation not permitted
$ docker run --rm --cap-drop ALL alpine:3.20 ping -c1 8.8.8.8
ping: permission denied (are you root?)
Root with no capabilities cannot chown, cannot open raw sockets (ping needs NET_RAW), and cannot bypass file permissions (DAC_OVERRIDE). Most services need none of the 14. Add back the specific one when something breaks: --cap-drop ALL --cap-add NET_BIND_SERVICE.
(NET_BIND_SERVICE normally allows binding ports below 1024. Docker sets net.ipv4.ip_unprivileged_port_start=0 inside containers, so in plain Docker any user can bind any port - other runtimes do not, so apps that listen on a high port like 8080 are the portable choice.)
A read-only root filesystem
$ docker run --rm --read-only alpine:3.20 touch /x
touch: /x: Read-only file system
$ docker run --rm --read-only --tmpfs /tmp alpine:3.20 touch /tmp/x
--read-only mounts the container's root filesystem read-only. An attacker who gets code execution cannot drop a binary or modify the app. Applications usually need a few writable paths - give them tmpfs or volumes precisely there:
$ docker run -d --read-only --name ro nginx:1.27
$ docker logs ro
... [emerg] 1#1: mkdir() "/var/cache/nginx/client_temp" failed (30: Read-only file system)
$ docker run -d --read-only --tmpfs /var/cache/nginx --tmpfs /run --name ro2 nginx:1.27
Finding those paths is the work: run it, read the error, add a tmpfs, repeat. docker diff (11.5) on a normal (writable) run lists exactly what it writes.
no-new-privileges
--security-opt no-new-privileges stops a process from gaining privileges through setuid binaries (sudo, su, passwd, 4.9) - you see NoNewPrivs: 1 in /proc/1/status. The same knob as NoNewPrivileges=yes in systemd (2.26).
--privileged: all of it, and more
--privileged gives every capability, switches off seccomp (a kernel filter that blocks dangerous system calls; Docker's default profile blocks dozens) and AppArmor (Ubuntu's rules for what files a program may touch), and gives access to all host devices. From a privileged container, root can mount the host's disk and chroot into it (make it its root directory) - full host root. Reserve it for tools that genuinely manage the host (the tonistiigi/binfmt installer in chapter 10) and never for an app.
The socket, again
docker run -v /var/run/docker.sock:/var/run/docker.sock docker:27-cli docker ps
/var/run/docker.sock is how you talk to dockerd, which runs as root (10.1). A container with the socket mounted can start a privileged container that mounts /. Mounting the socket = root on the host - common in CI build runners and "helper" containers, and worth flagging in every review.
User namespaces, briefly
Everything above is needed because, by default, uid 0 in a container is uid 0 on the host. A user namespace remaps it: Docker's userns-remap setting in daemon.json makes container root an unprivileged host uid, and rootless Docker or Podman (another container engine that needs no root daemon) run the whole engine as a normal user. Both cost compatibility (volume ownership gets interesting), which is why they are not the default.
A hardened run command
docker run -d --name api \
--user 10001:10001 \
--read-only --tmpfs /tmp \
--cap-drop ALL \
--security-opt no-new-privileges \
--memory 512m --cpus 1 --pids-limit 256 \
-p 127.0.0.1:8080:8080 \
api:2
Line by line: not root; nothing writable but /tmp; no capabilities; no setuid escalation; limits so it cannot starve the host (11.10); the port reachable only from this host (11.15).
Later (Ch 17): every line of this command becomes a field in a Kubernetes
securityContext- you will write the same thing again as YAML.
What you can now do
- read and decode a container's capabilities
- run a service with no capabilities, a read-only filesystem and no-new-privileges
- explain why
--privilegedand a mounted docker.sock mean host root