OnCallReady

Lesson 10.3 · Images & Builds · 21 min read

What a container actually is

In plain words

Picture a stack of clear plastic sheets, each with a few drawings on it. Look from the top and you see one picture made from all the sheets. The sheets from the shop are glued shut, so you cannot draw on them; to add something you lay one fresh sheet on top and draw there. Throw away the top sheet and your drawings are gone, but the shop sheets are untouched and can be shared with other kids.

Those glued sheets are image layers, read-only and shared by every container of that image. The fresh sheet is the container's writable layer, and overlayfs is what makes the stack look like one /. The kid drawing is an ordinary host process, fenced in by namespaces and cgroups. You can see all of it: ps, /proc/<pid>/ns, docker diff.

Why this matters

People call containers "lightweight VMs", and then get surprised: a file written in a container vanishes, free shows the wrong memory, an image built on a Mac will not start on a server. All of those make sense once you see what a container really is. You can prove it on this box with tools you already know.

What you need to know already: ps, PIDs and parent processes (Ch 3, "ps, and the STAT column"), /proc/<pid> (Ch 3), cgroup limits and /sys/fs/cgroup (Ch 2 "Hardening and cgroup limits", Ch 5 "cgroup OOM and exit 137"), mounts and symlinks (Ch 4), ss -tlnp (Ch 9).

The one-sentence answer

A container is a process on the host kernel (the core of Linux that runs every program), isolated with namespaces and limited with cgroups. That is the whole thing.

A VM (virtual machine, like the UTM VM this box imitates) is different: a hypervisor fakes a whole computer, and a second kernel boots inside it. A container has no hypervisor, no second kernel, no virtual hardware.

"Explain a container without saying lightweight VM" is a standard interview question, and the proof below is the answer.

Proof 1: it is in ps

$ docker run -d --name web nginx:1.27
3f4e1a2b5c6d7e8f...
$ ps -ef --forest | grep -A3 containerd-shim
root  2481     1  0 09:20 ?  00:00:00 /usr/bin/containerd-shim-runc-v2 -namespace moby -id 3f4e1a2b... -address /run/containerd/containerd.sock
root  2503  2481  0 09:20 ?  00:00:00  \_ nginx: master process nginx -g daemon off;
101   2541  2503  0 09:20 ?  00:00:00      \_ nginx: worker process
101   2542  2503  0 09:20 ?  00:00:00      \_ nginx: worker process

The flags: docker run -d starts the container detached (in the background) and prints its ID; --name web gives it a name to use instead of the ID. ps -ef --forest lists every process as a tree (Ch 3); grep -A3 shows the matching line and 3 lines after it.

nginx is an ordinary host process with a host PID (2503). Its parent is the shim, whose parent is PID 1 - systemd. The workers run as uid 101, which has no name on the host, so ps prints the number. docker top web shows the same thing from Docker's side, also with host PIDs:

$ docker top web
UID    PID    PPID   C   STIME   TTY   TIME       CMD
root   2503   2481   0   09:20   ?     00:00:00   nginx: master process nginx -g daemon off;
101    2541   2503   0   09:20   ?     00:00:00   nginx: worker process

The columns are the ps -ef ones: user, PID, parent PID, CPU, start time, terminal, CPU time, command.

Proof 2: namespaces are visible in /proc

Every process has one link per namespace in /proc/<pid>/ns. Two processes in the same namespace point at the same number (an inode number, Ch 4):

$ PID=$(docker inspect -f '{{.State.Pid}}' web)   # 2503 here
$ sudo readlink /proc/1/ns/net /proc/$PID/ns/net
net:[4026531840]
net:[4026532289]
$ sudo readlink /proc/1/ns/user /proc/$PID/ns/user
user:[4026531837]
user:[4026531837]

docker inspect prints everything Docker knows about a container as JSON; -f '{{.State.Pid}}' picks one field (the host PID of its main process). The {{ }} syntax is a Go template - Docker is written in Go, and this is Go's way of saying "print this field". readlink prints where a symlink points.

Different net: nginx has its own network interfaces and ports. Same user: Docker does not use user namespaces by default, so uid 0 in the container is uid 0 on the host - remember that for the security lesson. lsns lists namespaces (-t net: only network ones):

$ sudo lsns -t net
        NS TYPE NPROCS   PID USER COMMAND
4026531840 net     142     1 root /sbin/init
4026532289 net       3  2503 root nginx: master process nginx -g daemon off;

Columns: the namespace number, its type, how many processes are in it, the lowest PID in it, its owner, that process's command. The kinds of namespace:

pid     its own process tree - the first process is PID 1 inside
net     its own interfaces, routes, ports, iptables
mnt     its own mount table - the image becomes its /
uts     its own hostname
ipc     its own shared memory and semaphores
user    its own uid mapping (off by default in Docker)
cgroup  its own view of /sys/fs/cgroup

nsenter walks into a process's namespaces. -t $PID picks the target process, -n enters only its network namespace. You keep the host's filesystem and tools, but see the container's network - the trick for debugging images that have no shell:

$ sudo nsenter -t $PID -n ss -tlnp
State  Recv-Q Send-Q Local Address:Port  Peer Address:PortProcess
LISTEN 0      511    0.0.0.0:80          0.0.0.0:*    users:(("nginx",pid=1,fd=6))

The ss output is the one from Ch 9: nginx listens on port 80. Note pid=1: inside its own PID namespace, nginx is PID 1.

Proof 3: the limits are cgroups

$ docker run -d --name lim -m 200m nginx:1.27
$ cat /sys/fs/cgroup/system.slice/docker-$(docker inspect -f '{{.Id}}' lim).scope/memory.max
209715200
$ docker exec lim cat /sys/fs/cgroup/memory.max
209715200

-m 200m sets a 200 MiB memory limit. docker exec lim cat ... runs one more command inside the running container lim. The same file, seen from outside and inside: -m 200m did exactly what systemctl set-property demo MemoryMax=200M did in Ch 2.

Later (Ch 11): the next chapter goes deep on memory and CPU limits for containers.

Image vs container

An image = an ordered stack of read-only layers plus a config.

An image never changes; a "new version" is a new image with a new ID.

A container = an image + one writable layer on top + runtime settings (ports, mounts, limits, name). The stack is combined with overlayfs, a Linux filesystem that lays directories on top of each other and shows them as one:

upperdir   the container's writable layer     (deleted with the container)
lowerdir   the image layers, read-only, shared by every container of that image
merged     what the process sees as /

Everything a process writes goes to the upper dir (unless it writes into a volume, a directory stored outside the container - Ch 11). docker diff shows what a container changed:

$ docker exec web sh -c 'echo hi > /tmp/note'
$ docker diff web
C /run
A /run/nginx.pid
C /tmp
A /tmp/note
C /var
C /var/cache
C /var/cache/nginx
A /var/cache/nginx/client_temp
...

A added, C changed, D deleted. docker rm web (delete the container) throws all of it away - not "probably loses", destroys.

Why this is not a VM

VM          own kernel, virtual hardware, boots, minutes, GBs
container   the HOST kernel, no hardware, no boot, milliseconds, MBs

Consequences you will actually hit:

The three-sentence version for an interview

"A container is a normal Linux process that the kernel isolates with namespaces

same mechanism systemd uses for services. Its filesystem is an image: read-only layers stacked with overlayfs plus one writable layer that is discarded with the container. There is no guest kernel, which is why it starts in milliseconds and why it shares the host's kernel, architecture and security boundary."

What you can now do

Why it helps

This is the mental model behind half the tickets you will get. A container lost its uploaded files when it was recreated: they were in the writable layer, not a volume. An app inside a container sizes its memory use from free and gets OOM-killed: free shows the host, the cgroup file is the truth. Someone asks why a 'Ubuntu 22.04 container' has a 7.0 kernel: it is the host's. A security reviewer asks whether root in a container is root on the node: without user namespaces, yes. And when you debug a distroless image with no shell, nsenter -t <pid> -n lets you run the host's ss in its network namespace. It is also the first interview question for any container role, and proving it with ps and /proc is the answer that stands out.

Commands in this lesson

docker ps readlink lsns nsenter cat

FAQ

If a container is just a process, where is the isolation actually coming from?

From the kernel. Namespaces change what the process can see: its own PIDs, network interfaces, mount table, hostname, IPC. Cgroups limit what it can use: memory, CPU, PIDs, IO. Capabilities, seccomp and AppArmor restrict what it may do as root. Docker only configures these; runc applies them and gets out of the way. Remove any of them and the process sees more of the host.

Why does free inside a container show the host's memory?

free reads /proc/meminfo, which is not namespaced: it describes the whole machine. The container's limit is in its cgroup, /sys/fs/cgroup/memory.max from inside. Tools and runtimes that size themselves from meminfo get it wrong and can be OOM-killed. Modern JVMs read the cgroup limit; older ones and many scripts do not.

What is the difference between the host PID and the PID inside the container?

The container has its own PID namespace, so its first process sees itself as PID 1. The host sees the same process with a normal host PID, like 2503 in the lesson. docker top and ps on the host show host PIDs; ps inside the container shows namespace PIDs. Both refer to one process. You use host PIDs for nsenter, strace and /proc/<pid> on the host.

Where do files written inside a container go, and when are they lost?

Anything written outside a mounted volume goes to the container's writable overlay layer, the upperdir. It survives docker stop and docker start, because the container still exists. It is destroyed by docker rm, and most deploys replace the container with a new one, so treat it as gone on every deploy. docker diff shows what is there. Data that matters goes in a volume.

Can I run an amd64 image on this arm64 VM or my Mac?

Not natively. The container runs on the host CPU, and an amd64 binary fails with exec format error. With QEMU binfmt emulation installed it runs, slowly. Docker Desktop on macOS can also emulate. The real fix is multi-arch images, covered in the registry lesson. It matters because most servers are amd64 while your Mac is arm64.

In an interview Junior

Explain what a container is without saying "lightweight VM".

A container is a normal Linux process on the host kernel, isolated with namespaces and limited with cgroups.

There is no hypervisor and no guest kernel. So it starts in milliseconds, it is visible in the host's ps, uname -r shows the host kernel, an amd64 binary fails on arm64 with exec format error, and a kernel exploit is not contained.

Also asked: What is the difference between an image and a container? · Why does a file written inside a container disappear when you remove it? · Why does free inside a container show the host's memory?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.