Why this matters
Without limits, one container with a memory leak can take the whole host down with it. With limits set wrongly, a healthy Java service dies with exit 137 every few hours. Docker's limits are not new magic: they are the same cgroup files you set for systemd units in chapter 2.
What you need to know already: cgroups and
MemoryMax/CPUQuotafor systemd units (2.26); the OOM killer and the cgroup OOM that ends in 137 (5.7, 5.11); JVM heap vs the rest of its memory, and-XX:MaxRAMPercentage(5.13).
The same files as chapter 2
$ docker run -d --name box -m 512m --cpus 1.5 --pids-limit 200 nginx:1.27
$ docker exec box cat /sys/fs/cgroup/memory.max /sys/fs/cgroup/cpu.max /sys/fs/cgroup/pids.max
536870912
150000 100000
200
-m 512m(--memory) ->memory.max = 536870912bytes (512 x 1024 x 1024). Go over it and the kernel's cgroup OOM killer kills a process inside the cgroup.--cpus 1.5->cpu.max = 150000 100000: a quota of 150000 microseconds (150ms) of CPU time per period of 100000 microseconds (100ms). Use it up and the process is throttled - paused until the next period starts - not killed.--pids-limit 200->pids.max: at 200 processes and threads, creating another fails withEAGAIN("resource temporarily unavailable"). A fork bomb - a program that copies itself until the machine chokes - stays inside its box.
Inside the container its own cgroup appears as /sys/fs/cgroup (the cgroup namespace). From the host, the same values are under /sys/fs/cgroup/system.slice/docker-<full id>.scope/ - next to the systemd services you saw in 2.26. Without -m, memory.max is max: the container can use the whole host.
free lies, the cgroup does not
$ docker run --rm -m 512m alpine:3.20 free -m
total used free shared buff/cache available
Mem: 5782 1790 2428 6 1561 3816
Same columns as free in 5.3 - and that is the host's memory, not 512MB. free reads /proc/meminfo, which no namespace fences in. Anything that sizes itself from it - old JVMs, some Node and Python tools, worker counts based on nproc - believes it has the whole machine. The real limit is in memory.max.
OOM: what 137 looks like
# orders:slim = the image from chapter 10's multi-stage mission
docker run -d --name hog -m 256m -e JAVA_TOOL_OPTIONS=-Xmx300m orders:slim
docker ps -a --filter name=hog --format '{{.Status}}'
Exited (137) 5 seconds ago
docker inspect -f '{{.State.OOMKilled}}' hog
true
sudo dmesg | tail -2
Memory cgroup out of memory: Killed process 5120 (java) total-vm:3840212kB, anon-rss:241120kB, file-rss:21000kB, shmem-rss:0kB, UID:1000 pgtables:524kB oom_score_adj:0
OOMKilled: true plus the kernel's Memory cgroup out of memory line (you read these in 5.8) - the same cgroup OOM as chapter 5, not a host OOM: only processes in that container's cgroup were candidates.
The JVM reads the limit - and uses a quarter of it
The JVM (Java Virtual Machine) is the program that runs Java apps; its heap is where the app's objects live, capped by -Xmx (5.13). Since JDK 10 (-XX:+UseContainerSupport, on by default) the JVM reads the cgroup limit and sizes the heap from it: MaxRAMPercentage, default 25%.
$ docker run --rm -m 1g eclipse-temurin:21-jre java -XshowSettings:vm -version
VM settings:
Max. Heap Size (Estimated): 248.31M
Using VM: OpenJDK 64-Bit Server VM
$ docker run --rm -m 1g eclipse-temurin:21-jre java -XX:MaxRAMPercentage=75 -XshowSettings:vm -version
VM settings:
Max. Heap Size (Estimated): 744.94M
-XshowSettings:vm -version makes the JVM print the heap it picked and exit.
Two opposite failures follow:
- Heap too small: 25% of a 512MB container is 128MB of heap. A Java web service that needs 200MB throws
java.lang.OutOfMemoryError: Java heap spaceand exits 1 - the JVM's own limit; the container limit was never touched. - Heap too big:
-Xmxequal to the container limit. The heap fits, but the JVM also needs non-heap memory - metaspace (class data), thread stacks, code cache, direct buffers: 100-200MB for a typical service. Total > limit -> 137, OOMKilled.
The rule from 5.13, now for containers: limit = heap x 1.3 to 1.5, or set -XX:MaxRAMPercentage=75 (no -Xmx) and let the JVM do the arithmetic. JAVA_TOOL_OPTIONS is an environment variable the JVM reads extra flags from (it prints "Picked up JAVA_TOOL_OPTIONS: ..." on stderr) - handy when you cannot change the image's ENTRYPOINT: -e JAVA_TOOL_OPTIONS=-XX:MaxRAMPercentage=75.
CPU: throttling, not killing
# cpu-bound = a busy loop started with --cpus 0.5 (the limits mission)
docker stats --no-stream cpu-bound
CONTAINER ID NAME CPU % MEM USAGE / LIMIT MEM % ...
9a8b7c6d5e4f cpu-bound 50.02% 12.1MiB / 5.649GiB 0.21%
docker exec cpu-bound cat /sys/fs/cgroup/cpu.stat
usage_usec 18400211
user_usec 0
system_usec 0
nr_periods 3680
nr_throttled 3680
throttled_usec 176640000
--cpus 0.5: docker stats never goes above 50%. In cpu.stat: nr_periods = 100ms periods so far, nr_throttled = periods in which it hit its quota and was paused (here: every one), throttled_usec = total time spent paused. For a latency-sensitive service that shows up as small pauses every 100ms that nobody can explain - a CPU limit that is too tight hurts response times long before average CPU looks high.
Changing limits on a running container
docker update --memory 768m --memory-swap 768m api
docker update --cpus 2 api
docker update rewrites the cgroup files live, like systemctl set-property (2.26). Handy in an emergency; the real fix goes into the command that creates the container, or the next recreate undoes it. (--memory-swap is memory + swap together; when a container already has one, the memory limit cannot exceed it - update both.)
Later (Ch 17): Kubernetes sets these same cgroup files from a container's
resources; what you read here is exactly what it writes.
What you can now do
- read a container's memory, CPU and PID limits from inside and from the host
- tell a JVM heap OutOfMemoryError (exit 1) from a cgroup OOM kill (exit 137)
- size a container for a JVM, and spot CPU throttling in
cpu.stat