A pod restarts every few hours. kubectl describe pod says:
Last State: Terminated
Reason: OOMKilled
Exit Code: 137Read the number first
An exit code above 128 means "killed by a signal", and the signal is the code minus 128. 137 - 128 = 9 = SIGKILL. Nothing can catch SIGKILL, so the program had no chance to log anything. Something outside it decided it had to die. With memory, that something is the OOM killer. (The other usual sender is the kubelet itself, when a pod's grace period runs out.)
You can reproduce it on any systemd box:
$ sudo systemd-run --scope -p MemoryMax=200M memhog 400M
Killed
$ echo $?
137Host OOM vs cgroup OOM
- Host OOM: the whole machine ran out of memory. The kernel picks a victim by
oom_score(mostly: who uses the most), and it may not be the process that caused it. - cgroup OOM: one group of processes - a container, a pod, a systemd service - hit its own limit (
memory.max). The kill happens inside that group, even with plenty of free memory on the host. Container limits are this kind.
The kernel log says which one it was. journalctl -k | grep -i oom (or sudo dmesg -T) shows a Memory cgroup out of memory line for a cgroup OOM, with the cgroup named.
The JVM trap
The classic setup: a 1 GiB container limit and -Xmx1g. Looks tidy. It gets killed anyway, because the heap is not all the memory a JVM uses. On top of the heap come:
- metaspace (class metadata),
- thread stacks (one per thread),
- the JIT code cache,
- GC bookkeeping,
- direct/NIO buffers,
- the native libraries.
Heap = limit means the total is always over the limit. It's only a matter of when.
What to do:
- Leave headroom. Size the heap as a share of the limit and let the JVM do the maths:
-XX:MaxRAMPercentage=75with no-Xmx. Modern JVMs are container-aware: they read the cgroup limit instead of the host's RAM. - Measure the rest. Start with
-XX:NativeMemoryTracking=summary, thenjcmd <pid> VM.native_memory summaryshows where the non-heap memory goes. - Or raise the limit to fit heap + non-heap + a margin.
The diagnosis path, in order
kubectl describe pod/systemctl status: is it 137 / OOMKilled /oom-kill? A pod in CrashLoopBackOff starts with the same Last State check.- The kernel log: a cgroup OOM or a host OOM? A plain Docker container that keeps restarting reads the same way.
- The limit vs the real footprint (heap plus native), not just
-Xmx. - Fix the sizing, then watch memory over a few hours. A slow climb toward the limit is a leak, not a sizing problem.