OnCallReady

Lesson 11.29 · Docker Runtime & Networking · 12 min read

Container security: capabilities, read-only, and the socket

In plain words

Imagine a school janitor with a huge key ring that opens every door. For most jobs he needs only three keys. If he loses the ring, whoever finds it can enter anywhere. Better to give him just the keys for today's job, lock the supply room so nothing can be added, and never hand him the master key to the whole building.

Root in Linux is that key ring, split into about 40 capabilities. Docker gives a root container only 14 of them, and --cap-drop ALL takes away the rest. --read-only locks the filesystem, no-new-privileges stops setuid tricks, and --privileged or mounting /var/run/docker.sock is handing over the master key: root on the host.

Why this matters

All containers on a host share one kernel. If an attacker gets code running in a container, what they can do next depends on the flags you ran it with. A few flags turn "they own one app" into "they own nothing much" - and a few others turn it into "they own the host". This lesson is the checklist, with a reason for each item.

What you need to know already: root vs normal users and setuid binaries (4.9); systemd hardening like NoNewPrivileges= (2.26); /proc/PID/status (3.14); running images as non-root with USER and the docker group = root (10.40, 10.1).

Root in a container is not all of root

Linux splits root's powers into about 40 separate permissions called capabilities: CAP_CHOWN (change file owners), CAP_NET_RAW (raw network packets, used by ping), CAP_SYS_ADMIN (mounting and much else)... A process can hold some and not others. Root in a default container keeps only 14:

$ docker run --rm alpine:3.20 grep Cap /proc/1/status
CapInh:	0000000000000000
CapPrm:	00000000a80425fb
CapEff:	00000000a80425fb
CapBnd:	00000000a80425fb
CapAmb:	0000000000000000
$ capsh --decode=00000000a80425fb
0x00000000a80425fb=cap_chown,cap_dac_override,cap_fowner,cap_fsetid,cap_kill,cap_setgid,cap_setuid,cap_setpcap,cap_net_bind_service,cap_net_raw,cap_sys_chroot,cap_mknod,cap_audit_write,cap_setfcap

The kernel stores the sets as hex bitmasks, one bit per capability. The one that counts is CapEff (effective: what the process can use right now); CapBnd (bounding) is the most it could ever get. capsh --decode=HEX turns a mask into names (capsh is in the libcap2-bin package, on the host).

Missing on purpose: SYS_ADMIN (mount, most of "root"), NET_ADMIN (change interfaces and firewall rules), SYS_PTRACE (attach to other processes, like strace), SYS_MODULE (load kernel modules), SYS_TIME... So a root container cannot, say, load a kernel module.

A process running as a non-root user has an empty effective set (CapEff: 0000000000000000) whatever the bounding set says - another reason USER matters.

Drop everything, add back what you need

$ docker run --rm --cap-drop ALL alpine:3.20 grep CapEff /proc/1/status
CapEff:	0000000000000000
$ docker run --rm --cap-drop ALL alpine:3.20 chown 1000 /etc/hostname
chown: /etc/hostname: Operation not permitted
$ docker run --rm --cap-drop ALL alpine:3.20 ping -c1 8.8.8.8
ping: permission denied (are you root?)

Root with no capabilities cannot chown, cannot open raw sockets (ping needs NET_RAW), and cannot bypass file permissions (DAC_OVERRIDE). Most services need none of the 14. Add back the specific one when something breaks: --cap-drop ALL --cap-add NET_BIND_SERVICE.

(NET_BIND_SERVICE normally allows binding ports below 1024. Docker sets net.ipv4.ip_unprivileged_port_start=0 inside containers, so in plain Docker any user can bind any port - other runtimes do not, so apps that listen on a high port like 8080 are the portable choice.)

A read-only root filesystem

$ docker run --rm --read-only alpine:3.20 touch /x
touch: /x: Read-only file system
$ docker run --rm --read-only --tmpfs /tmp alpine:3.20 touch /tmp/x

--read-only mounts the container's root filesystem read-only. An attacker who gets code execution cannot drop a binary or modify the app. Applications usually need a few writable paths - give them tmpfs or volumes precisely there:

$ docker run -d --read-only --name ro nginx:1.27
$ docker logs ro
... [emerg] 1#1: mkdir() "/var/cache/nginx/client_temp" failed (30: Read-only file system)
$ docker run -d --read-only --tmpfs /var/cache/nginx --tmpfs /run --name ro2 nginx:1.27

Finding those paths is the work: run it, read the error, add a tmpfs, repeat. docker diff (11.5) on a normal (writable) run lists exactly what it writes.

no-new-privileges

--security-opt no-new-privileges stops a process from gaining privileges through setuid binaries (sudo, su, passwd, 4.9) - you see NoNewPrivs: 1 in /proc/1/status. The same knob as NoNewPrivileges=yes in systemd (2.26).

--privileged: all of it, and more

--privileged gives every capability, switches off seccomp (a kernel filter that blocks dangerous system calls; Docker's default profile blocks dozens) and AppArmor (Ubuntu's rules for what files a program may touch), and gives access to all host devices. From a privileged container, root can mount the host's disk and chroot into it (make it its root directory) - full host root. Reserve it for tools that genuinely manage the host (the tonistiigi/binfmt installer in chapter 10) and never for an app.

The socket, again

docker run -v /var/run/docker.sock:/var/run/docker.sock docker:27-cli docker ps

/var/run/docker.sock is how you talk to dockerd, which runs as root (10.1). A container with the socket mounted can start a privileged container that mounts /. Mounting the socket = root on the host - common in CI build runners and "helper" containers, and worth flagging in every review.

User namespaces, briefly

Everything above is needed because, by default, uid 0 in a container is uid 0 on the host. A user namespace remaps it: Docker's userns-remap setting in daemon.json makes container root an unprivileged host uid, and rootless Docker or Podman (another container engine that needs no root daemon) run the whole engine as a normal user. Both cost compatibility (volume ownership gets interesting), which is why they are not the default.

A hardened run command

docker run -d --name api \
  --user 10001:10001 \
  --read-only --tmpfs /tmp \
  --cap-drop ALL \
  --security-opt no-new-privileges \
  --memory 512m --cpus 1 --pids-limit 256 \
  -p 127.0.0.1:8080:8080 \
  api:2

Line by line: not root; nothing writable but /tmp; no capabilities; no setuid escalation; limits so it cannot starve the host (11.10); the port reachable only from this host (11.15).

Later (Ch 17): every line of this command becomes a field in a Kubernetes securityContext - you will write the same thing again as YAML.

What you can now do

Why it helps

Security reviews of Dockerfiles, compose files and run commands are routine platform work, and this lesson gives you the checklist: runs as non-root, drops capabilities, read-only root filesystem with tmpfs where needed, no-new-privileges, no --privileged, no docker.sock mounted into a CI runner or helper container. Each item has a clear reason you can explain in the PR comment.

It also explains breakage: an image that assumes root failing once someone runs it with --user, an entrypoint that switches users failing after --cap-drop ALL, or a hardened container crashing on Read-only file system until you add the right tmpfs. Knowing which control caused which error turns an afternoon of trial and error into one targeted --cap-add or --tmpfs.

Commands in this lesson

docker capsh

FAQ

If a container runs as root, is it root on the host?

By default it is uid 0 on the host, but with limits: only 14 capabilities, a default seccomp profile blocking risky syscalls, AppArmor confinement, and namespaces hiding the host's processes and network. That is much less than full root, but it is still the same uid 0, so if something breaks out, through a kernel bug, --privileged, a mounted docker.sock or a host path mount, it lands as real root. That is why non-root users and user namespaces matter.

What does --cap-drop ALL break?

Anything that needs a root privilege. Root with no capabilities cannot chown files, cannot bypass file permissions (DAC_OVERRIDE), cannot open raw sockets so ping fails without NET_RAW, and cannot change uid or gid, which breaks entrypoints that start as root and then switch user with gosu or su-exec. Most application servers need none of them. Add back the specific capability when something breaks, like --cap-add NET_BIND_SERVICE or --cap-add CHOWN.

Why does my read-only container fail at start?

Because the app writes somewhere it did not tell you about. nginx needs /var/cache/nginx and /run, many apps write to /tmp, JVMs write temp files and sometimes heap dumps. With --read-only those writes fail with Read-only file system. Mount a tmpfs or a volume exactly on those paths. docker diff on a normal run lists what the container writes, which is the quickest way to find them.

Why is mounting the Docker socket so dangerous?

/var/run/docker.sock is the Docker API with full control over the daemon, which runs as root. Anything that can talk to it can start a new container with --privileged and -v /:/host, then chroot into the host filesystem. So a container with the socket mounted, or a user in the docker group, is effectively root on the host. For CI builds, prefer rootless builders like BuildKit rootless or Kaniko-style builders, or isolated build VMs.

Does binding to port 80 need root or NET_BIND_SERVICE in a container?

In plain Docker, no: Docker sets net.ipv4.ip_unprivileged_port_start=0 inside each container's network namespace, so any user can bind any port, even with all capabilities dropped. On a normal Linux host, ports below 1024 need root or NET_BIND_SERVICE. Other container runtimes do not all set that sysctl, so the portable choice is to have the app listen on a high port like 8080 and publish it as 80 with -p 80:8080.

In an interview Junior

How would you harden a docker run command for a typical web service?

Each flag removes something an attacker could use:

docker run -d --name api \
  --user 10001:10001 \
  --read-only --tmpfs /tmp \
  --cap-drop ALL \
  --security-opt no-new-privileges \
  -m 512m --cpus 1 --pids-limit 200 \
  -p 127.0.0.1:8080:8080 \
  api:1

Never --privileged and never mount /var/run/docker.sock into an app - both mean root on the host.

Also asked: Why should containers not run as root? · What are Linux capabilities? · Why is mounting the Docker socket into a container dangerous?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.