Chapter 5 Memory & OOM
RSS vs VSZ vs PSS, why free memory is a lie, reading an OOM report, the OOM killer and its scores, host OOM vs cgroup OOM, and the Java program that sizes itself from the wrong number.
In plain words
Imagine a desk in a classroom. Your own books and notes (a program's own memory) take real space. The library keeps copies of books you read recently on a side shelf, just in case you want them again, and happily takes them away when someone needs room. When the desk is truly full and nothing can be moved, the teacher picks a pupil and clears their whole desk at once, usually the one with the biggest pile, not necessarily the one who caused the mess.
That is Linux memory. RSS is what is really on the desk, the page cache is the side shelf (counted in buff/cache, but available), and the OOM killer is the teacher. Some pupils also have their own fenced area with a limit (a cgroup with MemoryMax), and hitting that fence gets only them cleared, with exit code 137.
Why it matters on call
Memory is behind some of the most common and most misread production issues: services restarting with exit 137, Java services killed while their heap looks fine, alerts firing because "free" memory is low on a perfectly healthy box, a machine where the kernel killed the database instead of the leaking script. Knowing the difference between free and available, RSS and VSZ, host OOM and cgroup OOM, and heap and non-heap turns those into short investigations.
It is also the basis for sizing: setting MemoryMax and MemoryHigh, choosing MaxRAMPercentage for a Java service, deciding on swap, and protecting sshd so you can still log in. The three Notion questions for this topic (the exit 137 diagnosis path, why low free memory is normal, sizing a limit for a Java program) are standard SRE interview questions.
Lessons
- RSS, VSZ and shared
- free, available, and the page cache
- The OOM killer
- Reading an OOM report line by line
- cgroup OOM and exit 137
- JVM memory: what the heap is not
10 hands-on labs (missions, incidents and drills) run in the terminal: Open this chapter in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.
Questions people ask
Is it bad that Linux uses almost all my RAM?
Usually not. Linux uses otherwise idle memory as page cache for file contents, which makes reads faster and is released instantly when programs need memory. That shows up as low free and high buff/cache. Look at available in free -m: if it is healthy and there is no sustained swapping in vmstat, the box is fine. Unused RAM is wasted RAM.
What does exit code 137 mean for memory?
137 is 128 + 9: the process was killed by SIGKILL. For a service with a memory limit, the usual sender is the kernel's cgroup OOM killer after the process exceeded MemoryMax; systemd then reports Result: oom-kill. It can also come from a stop timeout or someone's kill -9. The kernel log (journalctl -k) with Memory cgroup out of memory confirms the OOM case.
Should servers have swap?
It depends. A modest amount of swap lets the kernel move rarely used pages out and absorbs short spikes, which is why Ubuntu creates /swap.img. Heavy swapping (thrashing, constant si/so in vmstat) makes a box unusably slow, often worse than an OOM kill. Some software, like the kubelet in the node incident, refuses to run with swap on. Databases often set a low swappiness rather than no swap.
Why does the JVM use more memory than -Xmx?
-Xmx limits only the Java heap. The process also needs metaspace for classes, a stack per thread, the JIT code cache, garbage collector structures, direct (off-heap) buffers and native allocations by libraries. Together those commonly add 200-500 MB or more. The kernel and a cgroup memory limit see the whole process, so a limit equal to -Xmx will eventually be exceeded.
What is the difference between MemoryHigh and MemoryMax?
MemoryMax is the hard limit: if the cgroup cannot be reclaimed below it, the kernel OOM-kills a process inside, exit 137. MemoryHigh is a soft limit: above it the kernel reclaims hard and slows the cgroup down, but never kills. Setting High a little below Max gives you a slow, visible service (the high counter in memory.events) before a dead one.