OnCallReady

Lesson 5.11 · Memory & OOM · 13 min read

cgroup OOM and exit 137

In plain words

Imagine a shared garden where each family has a fenced plot. If the whole garden floods, the gardener may have to clear any plot, starting with the biggest. But if one family overfills their own plot, only their plants get cut back, even though the rest of the garden has plenty of room.

The whole garden flooding is a host OOM. One plot overflowing its fence is a cgroup OOM: the fence is memory.max, set by MemoryMax= in a systemd unit. The killed process exits with 137 (128 + 9, SIGKILL). memory.high is a softer fence that slows the family down before cutting anything. The files under /sys/fs/cgroup/system.slice/orders.service/ are the plot's meter: memory.current, memory.peak, memory.events.

Two OOMs that look nothing alike in the log

A service keeps dying with exit code 137 while free shows gigabytes available. The machine is fine - the service hit its own memory limit. This lesson is the files that prove it and the arithmetic behind 137.

What you need to know already: 2.26 (cgroups, MemoryMax=, systemd-run), 5.7 (the OOM killer, host vs cgroup), 3.6 (signals), 1.7 (exit codes and $?).

Host OOM - the machine ran out. global_oom, CONSTRAINT_NONE. Everything on the box is a candidate. Usually means you under-provisioned, or something leaked.

cgroup OOM - one cgroup hit its own limit while the machine had plenty free. CONSTRAINT_MEMCG, and an oom_memcg= naming the cgroup. Only processes inside that cgroup are candidates.

For a systemd service, the second one shows up as Result: oom-kill, and the main process's exit code is 137.

Why 137

128 + signal number.  SIGKILL is 9.  128 + 9 = 137.

When a process is killed by a signal, the shell (and systemd) report its exit status as 128 plus the signal's number. Same arithmetic gives 143 for SIGTERM (128+15) - which is why SuccessExitStatus=143 appears in the units of Java services - and 130 for Ctrl+C (SIGINT, 2).

137 always means SIGKILL. With a memory limit in place, the sender is usually the cgroup OOM killer. The other common sender is systemd escalating to SIGKILL after a stop timeout (2.24) - which is why you confirm with the kernel log rather than assume.

sleep 100 & runs sleep in the background; kill -9 %1 SIGKILLs job 1; wait %1 waits for it and takes its exit status into $?:

$ sleep 100 &
[1] 17421
$ kill -9 %1; wait %1; echo $?
[1]+  Killed                  sleep 100
137

The files

Every cgroup is a directory under /sys/fs/cgroup/ (2.26), and its memory settings and counters are files in it:

$ cd /sys/fs/cgroup/system.slice/orders.service
$ cat memory.max memory.high memory.current memory.peak
max
max
626688000
626688000
$ cat memory.events
low 0
high 0
max 0
oom 0
oom_kill 0
oom_group_kill 0

A gotcha worth knowing: /sys/fs/cgroup/memory.max does not exist at the root of the hierarchy - the root cgroup (the whole machine) has no limit, so the file is absent. Limits live in the child cgroups: go to the unit's own path under system.slice.

Find any process's cgroup with cat /proc/<pid>/cgroup (0::/system.slice/orders.service) and prefix it with /sys/fs/cgroup.

Later (Ch 10): inside a container, the container's own cgroup appears as the root, so cat /sys/fs/cgroup/memory.max there is the container's limit. Runbooks written for one context confuse people in the other.

Setting them

sudo systemctl set-property orders MemoryHigh=900M MemoryMax=1200M   # live + persistent
sudo systemctl set-property --runtime orders MemoryMax=1200M          # until reboot
sudo systemctl edit orders                                            # [Service] MemoryMax=...

set-property (2.26) writes a drop-in under /etc/systemd/system.control/ and applies it to the running cgroup immediately - no restart. --runtime makes it temporary. Units accept K, M, G, a percentage of RAM (MemoryMax=25%) or infinity.

Trying it without writing a unit

$ sudo systemd-run --scope -p MemoryMax=200M memhog 400M
Running as unit: run-r2d4c6f0e1b2a4c89a1f4e3b2c1d0a9f8.scope
....
Killed
$ echo $?
137

systemd-run runs a command under systemd; --scope keeps it attached to your terminal but inside a new, temporary cgroup (a scope); -p MemoryMax=200M sets a property on that cgroup. memhog (a simulator tool) allocates 400M, the cgroup OOM killer fires at 200M, and your shell reports the SIGKILL as 137.

The same one-liner works for CPUQuota=, TasksMax= and IOWeight= - a very quick way to find out how a program behaves under a limit before you commit it to a unit file.

Later (Ch 16): in Kubernetes a container's memory limit becomes its cgroup's memory.max (exceed it: OOMKilled, exit 137), while its memory request is only used to decide which machine it runs on. When a whole machine runs low, Kubernetes also evicts (stops and moves) containers before the kernel's host OOM would act - a third way to die of memory, with its own signature.

What you can now do

Why it helps

Exit code 137 with Result: oom-kill is one of the most frequent service failures you will debug, and it is almost always this mechanism: the service exceeded its own limit, not the machine running out. Knowing where the cgroup files are, and that memory.events shows oom_kill independently of logs, gives you proof quickly.

It also clarifies sizing conversations: a hard limit protects the rest of the machine, a soft limit gives warning before the kill, and the limit is the thing to change when the machine was fine. systemd-run --scope -p MemoryMax= lets you test how a program behaves under a limit on any machine before you commit a value to a unit file.

Commands in this lesson

systemd-run sleep kill cat cd echo

FAQ

Why does /sys/fs/cgroup/memory.max not exist on my host?

You are looking at the root of the cgroup hierarchy, which is the whole machine, and the root cgroup has no memory limit, so the file is not present there. Limits live in child cgroups, for example /sys/fs/cgroup/system.slice/orders.service/memory.max. Use cat /proc/PID/cgroup to find a process's cgroup path and prefix it with /sys/fs/cgroup.

What is the difference between MemoryHigh and MemoryMax?

MemoryMax (memory.max) is the hard limit: when usage cannot be reclaimed below it, the cgroup OOM killer kills a process inside. MemoryHigh (memory.high) is a throttle: above it the kernel reclaims aggressively and slows the cgroup's allocations, but never kills. Setting High somewhat below Max gives the workload a slowdown and an observable signal (high events) before any kill.

Does page cache count against a service's memory limit?

Yes. Page cache for files a cgroup reads or writes is charged to that cgroup and counts in memory.current. The kernel reclaims it before OOM-killing, so clean cache rarely causes kills by itself, but dirty pages not yet written back and files in tmpfs (like /dev/shm) do count and cannot simply be dropped. memory.stat shows the file and anon split.

Is exit code 137 always an OOM kill?

No, it always means SIGKILL, and the OOM killer is only the most common sender for a service with a memory limit. A systemd stop that hit TimeoutStopSec and escalated to SIGKILL, or someone running kill -9, also produce 137. Confirm with Result: oom-kill in systemctl status, the oom_kill counter in memory.events, or the kernel log.

Can I change a service's memory limit without restarting it?

Yes. sudo systemctl set-property svc MemoryMax=1200M writes the new value into the running cgroup immediately and persists it as a drop-in under /etc/systemd/system.control/. Add --runtime to make it temporary. Lowering a limit below current usage triggers reclaim and possibly an immediate OOM kill, so check memory.current first.

In an interview Junior

What does exit code 137 mean for a service, and how do you confirm the cause?

137 = 128 + 9: the process was killed by SIGKILL. It does not say who sent it. With a memory limit it is usually the cgroup OOM killer - the service went over its MemoryMax while the machine may have had gigabytes free. The other common sender is systemd escalating to SIGKILL after a stop timeout.

To confirm:

Try it safely with sudo systemd-run --scope -p MemoryMax=200M memhog 400M; echo $? - it prints 137.

Also asked: What is the difference between MemoryMax and MemoryHigh? · How do you set and verify a memory limit on a running service? · Why does /sys/fs/cgroup/memory.max not exist on the host?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.