OnCallReady

Lesson 11.10 · Docker Runtime & Networking · 13 min read

Memory, CPU and PIDs: the cgroups again

In plain words

Imagine a shared kitchen where each cook gets a fridge shelf and a turn at the stove. If a cook tries to put more on their shelf than fits, the manager throws out some of their food: that is the memory limit, and it is harsh. If a cook wants the stove longer than their turn, they just have to wait for the next turn: that is the CPU limit, and it only slows them down.

Docker sets those budgets with cgroups, the same kernel feature systemd used in chapter 2. -m 512m writes memory.max, and going over it means the OOM killer, exit 137. --cpus 1.5 writes cpu.max, and going over it means throttling. --pids-limit writes pids.max. The files are right there under /sys/fs/cgroup.

Why this matters

Without limits, one container with a memory leak can take the whole host down with it. With limits set wrongly, a healthy Java service dies with exit 137 every few hours. Docker's limits are not new magic: they are the same cgroup files you set for systemd units in chapter 2.

What you need to know already: cgroups and MemoryMax/CPUQuota for systemd units (2.26); the OOM killer and the cgroup OOM that ends in 137 (5.7, 5.11); JVM heap vs the rest of its memory, and -XX:MaxRAMPercentage (5.13).

The same files as chapter 2

$ docker run -d --name box -m 512m --cpus 1.5 --pids-limit 200 nginx:1.27
$ docker exec box cat /sys/fs/cgroup/memory.max /sys/fs/cgroup/cpu.max /sys/fs/cgroup/pids.max
536870912
150000 100000
200

Inside the container its own cgroup appears as /sys/fs/cgroup (the cgroup namespace). From the host, the same values are under /sys/fs/cgroup/system.slice/docker-<full id>.scope/ - next to the systemd services you saw in 2.26. Without -m, memory.max is max: the container can use the whole host.

free lies, the cgroup does not

$ docker run --rm -m 512m alpine:3.20 free -m
               total        used        free      shared  buff/cache   available
Mem:            5782        1790        2428           6        1561        3816

Same columns as free in 5.3 - and that is the host's memory, not 512MB. free reads /proc/meminfo, which no namespace fences in. Anything that sizes itself from it - old JVMs, some Node and Python tools, worker counts based on nproc - believes it has the whole machine. The real limit is in memory.max.

OOM: what 137 looks like

# orders:slim = the image from chapter 10's multi-stage mission
docker run -d --name hog -m 256m -e JAVA_TOOL_OPTIONS=-Xmx300m orders:slim
docker ps -a --filter name=hog --format '{{.Status}}'
Exited (137) 5 seconds ago
docker inspect -f '{{.State.OOMKilled}}' hog
true
sudo dmesg | tail -2
Memory cgroup out of memory: Killed process 5120 (java) total-vm:3840212kB, anon-rss:241120kB, file-rss:21000kB, shmem-rss:0kB, UID:1000 pgtables:524kB oom_score_adj:0

OOMKilled: true plus the kernel's Memory cgroup out of memory line (you read these in 5.8) - the same cgroup OOM as chapter 5, not a host OOM: only processes in that container's cgroup were candidates.

The JVM reads the limit - and uses a quarter of it

The JVM (Java Virtual Machine) is the program that runs Java apps; its heap is where the app's objects live, capped by -Xmx (5.13). Since JDK 10 (-XX:+UseContainerSupport, on by default) the JVM reads the cgroup limit and sizes the heap from it: MaxRAMPercentage, default 25%.

$ docker run --rm -m 1g eclipse-temurin:21-jre java -XshowSettings:vm -version
VM settings:
    Max. Heap Size (Estimated): 248.31M
    Using VM: OpenJDK 64-Bit Server VM
$ docker run --rm -m 1g eclipse-temurin:21-jre java -XX:MaxRAMPercentage=75 -XshowSettings:vm -version
VM settings:
    Max. Heap Size (Estimated): 744.94M

-XshowSettings:vm -version makes the JVM print the heap it picked and exit.

Two opposite failures follow:

The rule from 5.13, now for containers: limit = heap x 1.3 to 1.5, or set -XX:MaxRAMPercentage=75 (no -Xmx) and let the JVM do the arithmetic. JAVA_TOOL_OPTIONS is an environment variable the JVM reads extra flags from (it prints "Picked up JAVA_TOOL_OPTIONS: ..." on stderr) - handy when you cannot change the image's ENTRYPOINT: -e JAVA_TOOL_OPTIONS=-XX:MaxRAMPercentage=75.

CPU: throttling, not killing

# cpu-bound = a busy loop started with --cpus 0.5 (the limits mission)
docker stats --no-stream cpu-bound
CONTAINER ID   NAME        CPU %     MEM USAGE / LIMIT     MEM %    ...
9a8b7c6d5e4f   cpu-bound   50.02%    12.1MiB / 5.649GiB    0.21%
docker exec cpu-bound cat /sys/fs/cgroup/cpu.stat
usage_usec 18400211
user_usec 0
system_usec 0
nr_periods 3680
nr_throttled 3680
throttled_usec 176640000

--cpus 0.5: docker stats never goes above 50%. In cpu.stat: nr_periods = 100ms periods so far, nr_throttled = periods in which it hit its quota and was paused (here: every one), throttled_usec = total time spent paused. For a latency-sensitive service that shows up as small pauses every 100ms that nobody can explain - a CPU limit that is too tight hurts response times long before average CPU looks high.

Changing limits on a running container

docker update --memory 768m --memory-swap 768m api
docker update --cpus 2 api

docker update rewrites the cgroup files live, like systemctl set-property (2.26). Handy in an emergency; the real fix goes into the command that creates the container, or the next recreate undoes it. (--memory-swap is memory + swap together; when a container already has one, the memory limit cannot exceed it - update both.)

Later (Ch 17): Kubernetes sets these same cgroup files from a container's resources; what you read here is exactly what it writes.

What you can now do

Why it helps

The single most common production container failure is exit 137 OOMKilled, and the most common cause is a JVM whose heap was sized without thinking about the container limit. Knowing that the JVM takes 25% of the limit by default, and that heap is not the whole footprint, lets you size orders and payments correctly: limit around heap times 1.3 to 1.5, or MaxRAMPercentage=75 without -Xmx.

CPU limits cause the other classic: latency spikes that nobody can explain, because the service is throttled every 100ms period. cpu.stat with a climbing nr_throttled proves it, and it is why teams argue about whether hard CPU limits are worth it at all. And "free says I have 6GB" inside a container is a trap that has misled many engineers and tools.

Commands in this lesson

docker

FAQ

Why does free inside the container show the host's memory?

Because /proc/meminfo, which free reads, is not namespaced: it always describes the whole host. The container's real limit is in the cgroup file /sys/fs/cgroup/memory.max, and its usage in memory.current. Any tool that sizes itself from /proc/meminfo or from the host's CPU count, such as old JVMs or worker counts based on nproc, believes it has the whole machine and can overrun its limit.

What happens when a container goes over its CPU limit?

Nothing dramatic: it is throttled, not killed. --cpus 1.5 becomes cpu.max = 150000 100000, meaning 150ms of CPU time per 100ms period across all its threads. Once it has used its quota, its threads are paused until the next period starts. You see it as nr_throttled and throttled_usec climbing in cpu.stat, and as latency spikes in the service, while docker stats never goes above the limit.

What is the difference between a host OOM and a cgroup OOM?

A host OOM happens when the whole machine runs out of memory: the kernel picks a victim from all processes by oom_score. A cgroup OOM happens when one cgroup reaches its memory.max: only processes inside that cgroup are candidates, even if the host has plenty free. Docker containers with -m hit the second kind. The kernel log says Memory cgroup out of memory and CONSTRAINT_MEMCG, and inspect shows OOMKilled true.

If the JVM reads the container limit, how can it still be OOMKilled?

The JVM sizes its heap from the limit, but the heap is only part of its memory. Metaspace, thread stacks, the JIT code cache, direct buffers and GC structures add 100-200MB or more for a typical service. If someone sets -Xmx equal to the container limit, heap plus the rest exceeds it and the kernel kills the process with 137. The opposite mistake, the default 25%, gives a heap too small and a Java OutOfMemoryError with exit 1.

Can I change a running container's limits?

Yes, docker update --memory 768m --memory-swap 768m api or docker update --cpus 2 api rewrites the cgroup files live, without a restart. It is useful in an emergency. But it only changes that container: the next recreate uses the original run command, or compose file again, so the real fix has to go there. If the container already has a swap limit, update both together or the change is refused.

In an interview Junior

How do you limit memory and CPU for a container, and what happens when it exceeds them?

They are the same cgroup files as a systemd unit's MemoryMax= / CPUQuota=:

docker stats shows usage against the limit; docker update changes it live (the real fix goes into the run command). Trap: free inside the container shows the host's memory; the truth is memory.max. For Java: the JVM takes 25% of the limit as heap by default; size the limit at heap x 1.3-1.5, or set -XX:MaxRAMPercentage=75.

Also asked: How would you size the memory limit for a Java service in a container? · What is the difference between a JVM OutOfMemoryError and a cgroup OOM kill? · What is CPU throttling?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.