Two OOMs that look nothing alike in the log
A service keeps dying with exit code 137 while free shows gigabytes available. The machine is fine - the service hit its own memory limit. This lesson is the files that prove it and the arithmetic behind 137.
What you need to know already: 2.26 (cgroups, MemoryMax=, systemd-run), 5.7 (the OOM killer, host vs cgroup), 3.6 (signals), 1.7 (exit codes and $?).
Host OOM - the machine ran out. global_oom, CONSTRAINT_NONE. Everything on the box is a candidate. Usually means you under-provisioned, or something leaked.
cgroup OOM - one cgroup hit its own limit while the machine had plenty free. CONSTRAINT_MEMCG, and an oom_memcg= naming the cgroup. Only processes inside that cgroup are candidates.
For a systemd service, the second one shows up as Result: oom-kill, and the main process's exit code is 137.
Why 137
128 + signal number. SIGKILL is 9. 128 + 9 = 137.
When a process is killed by a signal, the shell (and systemd) report its exit status as 128 plus the signal's number. Same arithmetic gives 143 for SIGTERM (128+15) - which is why SuccessExitStatus=143 appears in the units of Java services - and 130 for Ctrl+C (SIGINT, 2).
137 always means SIGKILL. With a memory limit in place, the sender is usually the cgroup OOM killer. The other common sender is systemd escalating to SIGKILL after a stop timeout (2.24) - which is why you confirm with the kernel log rather than assume.
sleep 100 & runs sleep in the background; kill -9 %1 SIGKILLs job 1; wait %1 waits for it and takes its exit status into $?:
$ sleep 100 &
[1] 17421
$ kill -9 %1; wait %1; echo $?
[1]+ Killed sleep 100
137
The files
Every cgroup is a directory under /sys/fs/cgroup/ (2.26), and its memory settings and counters are files in it:
$ cd /sys/fs/cgroup/system.slice/orders.service
$ cat memory.max memory.high memory.current memory.peak
max
max
626688000
626688000
$ cat memory.events
low 0
high 0
max 0
oom 0
oom_kill 0
oom_group_kill 0
memory.max- the hard limit (MemoryMax=).maxmeans none. Exceed it and reclaim fails: the cgroup OOM killer runs inside the cgroup.memory.high- the soft limit (MemoryHigh=). Above it the kernel slows the cgroup down and reclaims its memory aggressively: the processes get slow, nothing is killed. It is the early warning a hard limit does not give you.memory.current/memory.peak- bytes now, and the highest since the cgroup was created.memory.events- counters:hightimes over memory.high,maxtimes it hit memory.max,oom_killprocesses killed. A non-zerooom_killis proof, independent of any log.memory.stat- the breakdown:anon,file,kernel,shmem, ...
A gotcha worth knowing: /sys/fs/cgroup/memory.max does not exist at the root of the hierarchy - the root cgroup (the whole machine) has no limit, so the file is absent. Limits live in the child cgroups: go to the unit's own path under system.slice.
Find any process's cgroup with cat /proc/<pid>/cgroup (0::/system.slice/orders.service) and prefix it with /sys/fs/cgroup.
Later (Ch 10): inside a container, the container's own cgroup appears as the root, so
cat /sys/fs/cgroup/memory.maxthere is the container's limit. Runbooks written for one context confuse people in the other.
Setting them
sudo systemctl set-property orders MemoryHigh=900M MemoryMax=1200M # live + persistent
sudo systemctl set-property --runtime orders MemoryMax=1200M # until reboot
sudo systemctl edit orders # [Service] MemoryMax=...
set-property (2.26) writes a drop-in under /etc/systemd/system.control/ and applies it to the running cgroup immediately - no restart. --runtime makes it temporary. Units accept K, M, G, a percentage of RAM (MemoryMax=25%) or infinity.
Trying it without writing a unit
$ sudo systemd-run --scope -p MemoryMax=200M memhog 400M
Running as unit: run-r2d4c6f0e1b2a4c89a1f4e3b2c1d0a9f8.scope
....
Killed
$ echo $?
137
systemd-run runs a command under systemd; --scope keeps it attached to your terminal but inside a new, temporary cgroup (a scope); -p MemoryMax=200M sets a property on that cgroup. memhog (a simulator tool) allocates 400M, the cgroup OOM killer fires at 200M, and your shell reports the SIGKILL as 137.
The same one-liner works for CPUQuota=, TasksMax= and IOWeight= - a very quick way to find out how a program behaves under a limit before you commit it to a unit file.
Later (Ch 16): in Kubernetes a container's memory limit becomes its cgroup's
memory.max(exceed it:OOMKilled, exit 137), while its memory request is only used to decide which machine it runs on. When a whole machine runs low, Kubernetes also evicts (stops and moves) containers before the kernel's host OOM would act - a third way to die of memory, with its own signature.
What you can now do
- Decode 137 (and 143, 130) as 128 + a signal number.
- Prove a cgroup OOM from
memory.eventsand the kernel log. - Try any command under a memory limit with
systemd-run --scope -p MemoryMax=.