OnCallReady

MemoryKubernetesJVM · 2 min read

Exit code 137, OOMKilled, and the JVM that fits its heap but not its container

137 = 128 + SIGKILL. How to tell a cgroup OOM kill from a host OOM, why -Xmx equal to the memory limit gets your Java pod killed, and how to size it.

A pod restarts every few hours. kubectl describe pod says:

output
    Last State:     Terminated
      Reason:       OOMKilled
      Exit Code:    137

Read the number first

An exit code above 128 means "killed by a signal", and the signal is the code minus 128. 137 - 128 = 9 = SIGKILL. Nothing can catch SIGKILL, so the program had no chance to log anything. Something outside it decided it had to die. With memory, that something is the OOM killer. (The other usual sender is the kubelet itself, when a pod's grace period runs out.)

You can reproduce it on any systemd box:

terminal
$ sudo systemd-run --scope -p MemoryMax=200M memhog 400M
Killed
$ echo $?
137

Host OOM vs cgroup OOM

  • Host OOM: the whole machine ran out of memory. The kernel picks a victim by oom_score (mostly: who uses the most), and it may not be the process that caused it.
  • cgroup OOM: one group of processes - a container, a pod, a systemd service - hit its own limit (memory.max). The kill happens inside that group, even with plenty of free memory on the host. Container limits are this kind.

The kernel log says which one it was. journalctl -k | grep -i oom (or sudo dmesg -T) shows a Memory cgroup out of memory line for a cgroup OOM, with the cgroup named.

The JVM trap

The classic setup: a 1 GiB container limit and -Xmx1g. Looks tidy. It gets killed anyway, because the heap is not all the memory a JVM uses. On top of the heap come:

  • metaspace (class metadata),
  • thread stacks (one per thread),
  • the JIT code cache,
  • GC bookkeeping,
  • direct/NIO buffers,
  • the native libraries.

Heap = limit means the total is always over the limit. It's only a matter of when.

What to do:

  • Leave headroom. Size the heap as a share of the limit and let the JVM do the maths: -XX:MaxRAMPercentage=75 with no -Xmx. Modern JVMs are container-aware: they read the cgroup limit instead of the host's RAM.
  • Measure the rest. Start with -XX:NativeMemoryTracking=summary, then jcmd <pid> VM.native_memory summary shows where the non-heap memory goes.
  • Or raise the limit to fit heap + non-heap + a margin.

The diagnosis path, in order

  1. kubectl describe pod / systemctl status: is it 137 / OOMKilled / oom-kill? A pod in CrashLoopBackOff starts with the same Last State check.
  2. The kernel log: a cgroup OOM or a host OOM? A plain Docker container that keeps restarting reads the same way.
  3. The limit vs the real footprint (heap plus native), not just -Xmx.
  4. Fix the sizing, then watch memory over a few hours. A slow climb toward the limit is a leak, not a sizing problem.

OnCallReady is free, with no ads and no tracking. RSS · All posts