Why this matters
People call containers "lightweight VMs", and then get surprised: a file written in a container vanishes, free shows the wrong memory, an image built on a Mac will not start on a server. All of those make sense once you see what a container really is. You can prove it on this box with tools you already know.
What you need to know already: ps, PIDs and parent processes (Ch 3, "ps, and the STAT column"), /proc/<pid> (Ch 3), cgroup limits and /sys/fs/cgroup (Ch 2 "Hardening and cgroup limits", Ch 5 "cgroup OOM and exit 137"), mounts and symlinks (Ch 4), ss -tlnp (Ch 9).
The one-sentence answer
A container is a process on the host kernel (the core of Linux that runs every program), isolated with namespaces and limited with cgroups. That is the whole thing.
- A namespace gives a process its own private view of one kind of system resource: its own list of processes, its own network interfaces, its own mounts. It still runs on the same kernel as everything else.
- A cgroup caps what a group of processes may use (memory, CPU) - the same mechanism as
MemoryMax=in a systemd unit.
A VM (virtual machine, like the UTM VM this box imitates) is different: a hypervisor fakes a whole computer, and a second kernel boots inside it. A container has no hypervisor, no second kernel, no virtual hardware.
"Explain a container without saying lightweight VM" is a standard interview question, and the proof below is the answer.
Proof 1: it is in ps
$ docker run -d --name web nginx:1.27
3f4e1a2b5c6d7e8f...
$ ps -ef --forest | grep -A3 containerd-shim
root 2481 1 0 09:20 ? 00:00:00 /usr/bin/containerd-shim-runc-v2 -namespace moby -id 3f4e1a2b... -address /run/containerd/containerd.sock
root 2503 2481 0 09:20 ? 00:00:00 \_ nginx: master process nginx -g daemon off;
101 2541 2503 0 09:20 ? 00:00:00 \_ nginx: worker process
101 2542 2503 0 09:20 ? 00:00:00 \_ nginx: worker process
The flags: docker run -d starts the container detached (in the background) and prints its ID; --name web gives it a name to use instead of the ID. ps -ef --forest lists every process as a tree (Ch 3); grep -A3 shows the matching line and 3 lines after it.
nginx is an ordinary host process with a host PID (2503). Its parent is the shim, whose parent is PID 1 - systemd. The workers run as uid 101, which has no name on the host, so ps prints the number. docker top web shows the same thing from Docker's side, also with host PIDs:
$ docker top web
UID PID PPID C STIME TTY TIME CMD
root 2503 2481 0 09:20 ? 00:00:00 nginx: master process nginx -g daemon off;
101 2541 2503 0 09:20 ? 00:00:00 nginx: worker process
The columns are the ps -ef ones: user, PID, parent PID, CPU, start time, terminal, CPU time, command.
Proof 2: namespaces are visible in /proc
Every process has one link per namespace in /proc/<pid>/ns. Two processes in the same namespace point at the same number (an inode number, Ch 4):
$ PID=$(docker inspect -f '{{.State.Pid}}' web) # 2503 here
$ sudo readlink /proc/1/ns/net /proc/$PID/ns/net
net:[4026531840]
net:[4026532289]
$ sudo readlink /proc/1/ns/user /proc/$PID/ns/user
user:[4026531837]
user:[4026531837]
docker inspect prints everything Docker knows about a container as JSON; -f '{{.State.Pid}}' picks one field (the host PID of its main process). The {{ }} syntax is a Go template - Docker is written in Go, and this is Go's way of saying "print this field". readlink prints where a symlink points.
Different net: nginx has its own network interfaces and ports. Same user: Docker does not use user namespaces by default, so uid 0 in the container is uid 0 on the host - remember that for the security lesson. lsns lists namespaces (-t net: only network ones):
$ sudo lsns -t net
NS TYPE NPROCS PID USER COMMAND
4026531840 net 142 1 root /sbin/init
4026532289 net 3 2503 root nginx: master process nginx -g daemon off;
Columns: the namespace number, its type, how many processes are in it, the lowest PID in it, its owner, that process's command. The kinds of namespace:
pid its own process tree - the first process is PID 1 inside
net its own interfaces, routes, ports, iptables
mnt its own mount table - the image becomes its /
uts its own hostname
ipc its own shared memory and semaphores
user its own uid mapping (off by default in Docker)
cgroup its own view of /sys/fs/cgroup
nsenter walks into a process's namespaces. -t $PID picks the target process, -n enters only its network namespace. You keep the host's filesystem and tools, but see the container's network - the trick for debugging images that have no shell:
$ sudo nsenter -t $PID -n ss -tlnp
State Recv-Q Send-Q Local Address:Port Peer Address:PortProcess
LISTEN 0 511 0.0.0.0:80 0.0.0.0:* users:(("nginx",pid=1,fd=6))
The ss output is the one from Ch 9: nginx listens on port 80. Note pid=1: inside its own PID namespace, nginx is PID 1.
Proof 3: the limits are cgroups
$ docker run -d --name lim -m 200m nginx:1.27
$ cat /sys/fs/cgroup/system.slice/docker-$(docker inspect -f '{{.Id}}' lim).scope/memory.max
209715200
$ docker exec lim cat /sys/fs/cgroup/memory.max
209715200
-m 200m sets a 200 MiB memory limit. docker exec lim cat ... runs one more command inside the running container lim. The same file, seen from outside and inside: -m 200m did exactly what systemctl set-property demo MemoryMax=200M did in Ch 2.
Later (Ch 11): the next chapter goes deep on memory and CPU limits for containers.
Image vs container
An image = an ordered stack of read-only layers plus a config.
- A layer is a bundle (a tar archive) of the files one build step added or changed.
- The config is a small JSON file: what to run, environment variables, the user, the working directory, the ports it listens on.
An image never changes; a "new version" is a new image with a new ID.
A container = an image + one writable layer on top + runtime settings (ports, mounts, limits, name). The stack is combined with overlayfs, a Linux filesystem that lays directories on top of each other and shows them as one:
upperdir the container's writable layer (deleted with the container)
lowerdir the image layers, read-only, shared by every container of that image
merged what the process sees as /
Everything a process writes goes to the upper dir (unless it writes into a volume, a directory stored outside the container - Ch 11). docker diff shows what a container changed:
$ docker exec web sh -c 'echo hi > /tmp/note'
$ docker diff web
C /run
A /run/nginx.pid
C /tmp
A /tmp/note
C /var
C /var/cache
C /var/cache/nginx
A /var/cache/nginx/client_temp
...
A added, C changed, D deleted. docker rm web (delete the container) throws all of it away - not "probably loses", destroys.
Why this is not a VM
VM own kernel, virtual hardware, boots, minutes, GBs
container the HOST kernel, no hardware, no boot, milliseconds, MBs
Consequences you will actually hit:
uname -rin any container prints the host's kernel (7.0.0-31-generic). An "Ubuntu 22.04 container" is 22.04's files and programs (its userland) running on this kernel.- A kernel bug or a kernel exploit is not contained by Docker.
- The CPU type is not emulated. This box is arm64 (the chip family in Apple Silicon Macs,
uname -msays aarch64, Ch 1); most servers are amd64 (Intel/AMD). A program compiled for amd64 fails here withexec format error. That is the single most common image problem on Apple Silicon teams. freeinside a container shows the host's memory, not the limit (/proc/meminfois not namespaced). Tools that size themselves from it get it wrong; the cgroup filememory.maxis the truth (Ch 5).
The three-sentence version for an interview
"A container is a normal Linux process that the kernel isolates with namespaces
- its own PID tree, network, mounts and hostname - and limits with cgroups, the
same mechanism systemd uses for services. Its filesystem is an image: read-only layers stacked with overlayfs plus one writable layer that is discarded with the container. There is no guest kernel, which is why it starts in milliseconds and why it shares the host's kernel, architecture and security boundary."
What you can now do
- Find a container's processes in
psand its namespaces in/proc/<pid>/ns. - Use
nsenter -nto run host tools inside a container's network. - Explain image vs container, and why files written in a container disappear.