OnCallReady

Lesson 5.13 · Memory & OOM · 15 min read

JVM memory: what the heap is not

In plain words

Imagine packing for a trip with a suitcase limit of 20 kilos. You carefully weigh your clothes bag at exactly 20 kilos and call it done. At the airport the whole suitcase is weighed, including the shoes, the wash bag, the charger and the book, and it is over the limit.

The Java heap is the clothes bag. -Xmx1g limits only that. The JVM process also carries metaspace (class data), a stack for every thread, the JIT code cache, garbage collector bookkeeping, direct buffers and native libraries. MemoryMax weighs the whole suitcase. Set the limit equal to -Xmx and it will be exceeded. Better: set the limit, drop -Xmx, and let -XX:MaxRAMPercentage=75 size the heap from it, leaving room for everything else.

The heap is one region of several

Two services on this box are Java programs (orders and payments). A Java service with a 1 GB memory limit and "a 1 GB heap" gets OOM-killed over and over, and its own log shows no error at all. This lesson is where a Java program's memory actually goes, so you can give it a limit it survives.

What you need to know already: 5.1 (RSS, what the JVM is), 5.11 (cgroup limits, memory.max, exit 137), 2.3 (drop-ins with systemctl edit).

Recap from 5.1: a Java program runs inside the JVM, the java process. The program's data (its objects) live in the heap, a region the JVM manages. The JVM's garbage collector (GC) walks the heap from time to time and frees objects nothing uses any more - Java programs never free memory themselves.

-Xmx is a java command-line option: the maximum heap size (-Xmx1g = 1 GiB). The catch: the heap is only one of the JVM's memory regions.

-Xmx / MaxRAMPercentage   Java heap: your objects. The only part -Xmx limits.
-XX:MaxMetaspaceSize      metaspace: class metadata. Unlimited by default; grows
                          with the number of classes (a typical web service:
                          80-150 MB).
-Xss                      a stack per thread. Reserved: 1 MiB on x86-64
                          (2 MiB on arm64); only the touched pages count as RSS.
                          200 threads = up to a few hundred MB.
-XX:ReservedCodeCacheSize the JIT's compiled code. 240 MB reserved by default,
                          typically 30-80 MB used.
-XX:MaxDirectMemorySize   direct buffers: memory for network and file I/O
                          outside the heap. Defaults to the max heap size.
                          Invisible to heap dumps and to GC logs.
GC structures             the garbage collector's bookkeeping: a few % of the heap.
malloc / native libs      zlib, SSL - whatever C code inside the JVM allocates.

The words in that table:

A process's RSS is the sum. That is why a JVM with -Xmx1g sits at 1.3-1.5 GB of RSS once warmed up, and why limit = -Xmx is an OOM kill waiting for the first busy minute.

What the JVM decides by itself

java -XX:+PrintFlagsFinal -version prints every JVM setting with the value it would use on this machine (-version so it exits straight away); grep keeps the memory ones:

$ java -XX:+PrintFlagsFinal -version | grep -E 'UseContainerSupport|MaxRAMPercentage|InitialRAMPercentage|MaxHeapSize'
     bool UseContainerSupport         = true
   double MaxRAMPercentage            = 25.000000
   double InitialRAMPercentage        = 1.562500
   size_t MaxHeapSize                 = 1610612736

Measuring instead of guessing: Native Memory Tracking

Start the JVM with -XX:NativeMemoryTracking=summary (NMT; a few % overhead) and ask it with jcmd <pid> VM.native_memory summary (jcmd sends commands to a running JVM):

# on a machine with the full JDK, JVM started with -XX:NativeMemoryTracking=summary
jcmd 1210 VM.native_memory summary
1210:

Native Memory Tracking:

Total: reserved=2711393KB, committed=731245KB
-                 Java Heap (reserved=524288KB, committed=524288KB)
-                     Class (reserved=1063173KB, committed=16453KB)
-                    Thread (reserved=66660KB, committed=66660KB)
                            (thread #64)
-                      Code (reserved=248530KB, committed=38978KB)
-                        GC (reserved=58374KB, committed=58374KB)
-                  Internal (reserved=1285KB, committed=1285KB)
-                     Other (reserved=16452KB, committed=16452KB)

reserved is address space (VSZ-like); committed is memory actually backed, and that is what counts against the limit. Class = metaspace, Thread = stacks, Code = the code cache. Heap 512 MiB, plus ~200 MiB of everything else: that is the real footprint of this app, and the number to size the limit from. (jcmd ships with the full JDK, not with every Java install.)

Later (Ch 20): jcmd, GC logs and heap dumps in depth.

Sizing, in order of preference

  1. Let the heap follow the limit. Set MemoryMax= on the unit, drop -Xmx, and use -XX:MaxRAMPercentage=75 (or 70-80). Change the limit later and the heap follows; the two can never be set to the same number by accident.
  2. Fixed heap, derived limit. If the heap must be explicit: limit = heap x 1.3-1.5, or better, heap + measured non-heap (NMT) + 10-20%.
  3. Watch the right number. memory.current of the unit's cgroup under real load, and how full the heap actually gets. A heap that never passes 40% is a limit you can lower.

To change the java command line of a unit in a drop-in, clear ExecStart= first, then set it again (otherwise the unit ends up with two ExecStart= lines, which systemd refuses for a normal service - 2.3):

[Service]
ExecStart=
ExecStart=/usr/bin/java -XX:MaxRAMPercentage=75 -jar /opt/app/payments.jar

The failure modes and how each one looks

exit 137, CONSTRAINT_MEMCG, heap fine          limit too small for heap + non-heap
java.lang.OutOfMemoryError: Java heap space    heap too small (or a leak) - the JVM
                                               itself refused, the process may live on
OutOfMemoryError: Metaspace                    MaxMetaspaceSize, or ever more classes
OutOfMemoryError: Direct buffer memory         MaxDirectMemorySize
OutOfMemoryError: unable to create native      thread limit (TasksMax / pids.max),
  thread                                       not memory at all

OutOfMemoryError is the JVM's own error: it refused an allocation in one of its regions and says so, with a stack trace (the list of functions it was in) in the application log.

The first line is the kernel's decision and leaves nothing in the application log - the process is simply gone. The others are the JVM's decisions and leave a stack trace. Which log has the evidence tells you which kind you have.

Add -XX:+ExitOnOutOfMemoryError so a heap OOM kills the process (and systemd's Restart= starts a fresh one) instead of leaving a half-dead JVM serving errors, and -XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=/dumps if you have somewhere big enough to put a heap-sized file.

Later (Ch 10, Ch 16): this is the "JVM in a container" trap: a container's memory limit is exactly this cgroup limit, and UseContainerSupport is named after it.

What you can now do

Why it helps

A Java program whose heap is set to the same number as its memory limit is one of the most common causes of services dying with exit 137, and "how would you size a memory limit for a JVM" is a standard interview question. The symptom is confusing: heap numbers look healthy, there is no OutOfMemoryError in the logs, and the process just dies, because the kernel, not the JVM, made the decision.

Knowing the non-heap regions, UseContainerSupport, MaxRAMPercentage and Native Memory Tracking lets you fix it with numbers instead of doubling limits by guesswork. Reading which log has the evidence (kernel log versus application stack trace) tells you whether the heap, metaspace, direct memory, threads or the limit itself is the problem.

Commands in this lesson

java

FAQ

Should I use -Xmx or MaxRAMPercentage under a memory limit?

Prefer MaxRAMPercentage (for example 70-80) without -Xmx. The JVM reads its cgroup limit and sizes the heap as a percentage of it, so changing the limit automatically resizes the heap and the two cannot drift apart. Use an explicit -Xmx only when you need a fixed heap, and then derive the limit from it with room for non-heap memory. Setting both means -Xmx wins.

Why is the default heap only 25% of the memory limit?

MaxRAMPercentage defaults to 25, a conservative value chosen for machines running many processes, where the JVM should not grab most of RAM. When the JVM is the only significant process under its limit, 25% wastes most of it: a 4 GiB limit gets a 1 GiB heap. Raise it deliberately, usually to 70-80%, leaving the rest for metaspace, threads, code cache and direct buffers.

How do I see non-heap memory usage?

Start the JVM with -XX:NativeMemoryTracking=summary (small overhead) and run jcmd PID VM.native_memory summary: it shows reserved and committed memory for heap, class (metaspace), thread, code, GC, internal and other categories. Committed is what counts against the limit. Compare the total with the process RSS and the cgroup's memory.current; the difference is native allocations outside the JVM's tracking.

What is the difference between OutOfMemoryError and an OOM kill?

java.lang.OutOfMemoryError is the JVM refusing an allocation inside one of its own regions (heap, metaspace, direct memory) or failing to create a thread. It leaves a stack trace, and the process may keep running. An OOM kill is the kernel killing the whole process because the cgroup limit was exceeded; there is no Java stack trace, just exit 137 and Result: oom-kill. They need different fixes.

What do ExitOnOutOfMemoryError and HeapDumpOnOutOfMemoryError do?

-XX:+ExitOnOutOfMemoryError makes the JVM exit on the first OutOfMemoryError instead of continuing in a broken state, so systemd's Restart= starts it again cleanly. -XX:+HeapDumpOnOutOfMemoryError with -XX:HeapDumpPath=/dumps writes a heap dump for analysis. The dump is as large as the heap, so it needs a disk with enough space, and writing it takes time before exit.

In an interview Junior

How would you size a memory limit for a Java service?

First, know that the heap is only one part. A JVM's RSS is the heap (what -Xmx limits) plus metaspace, a stack per thread, the code cache, direct buffers, GC bookkeeping and native libraries. A JVM with -Xmx1g sits at 1.3-1.5 GB of RSS, so a limit equal to -Xmx is an OOM kill waiting for the first busy minute.

Then, in order of preference:

  1. Let the heap follow the limit: set MemoryMax= on the unit, drop -Xmx, use -XX:MaxRAMPercentage=75. UseContainerSupport makes the JVM read its cgroup limit as "the RAM".
  2. Fixed heap, derived limit: limit = heap x 1.3-1.5, or better heap + measured non-heap (Native Memory Tracking, jcmd PID VM.native_memory summary) + 10-20%.
  3. Watch the unit's memory.current under real load.

How each failure looks: exit 137 with nothing in the app log = the kernel (limit too small); OutOfMemoryError: Java heap space in the log = the heap itself.

Also asked: A Java service is OOM-killed but its heap never goes above 60%. What is happening? · What does UseContainerSupport do? · Without -Xmx, how much heap does the JVM pick?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.