OnCallReady

Lesson 20.8 · JVM Internals · 13 min read

Native memory tracking: where the rest of the RSS goes

In plain words

Imagine your parents give you pocket money and ask where it went. You only remember the big things: the comic and the cinema ticket. The rest just disappeared. Next week you keep a little notebook and write down every purchase, sorted into snacks, toys and bus fares. Now you can see that bus fares doubled.

Native Memory Tracking is the JVM keeping that notebook for the memory outside the heap. You have to start it from the beginning with -XX:NativeMemoryTracking=summary, because the notebook cannot be filled in afterwards. Then jcmd PID VM.native_memory summary shows each category, and baseline plus summary.diff shows which one grew, like Thread going from 56 to 422 threads while the heap stayed flat.

Turning it on

The problem. The heap is fine, RSS keeps growing, and the pod is OOMKilled anyway. Something outside the heap grows - threads, metaspace, buffers - and only the JVM can tell you which. NMT (Native Memory Tracking) is that report.

What you need to know already: the memory areas of a JVM (20.1), jcmd and attaching (20.3), systemd drop-ins (2.3), cgroup OOM (5.11).

NMT is the JVM accounting for its own native allocations. It must be enabled at startup - there is no attaching it later:

-XX:NativeMemoryTracking=summary     per category (5-10% overhead in the worst case, usually ~1-2%)
-XX:NativeMemoryTracking=detail      per call site - for the JVM developers, rarely needed

Without it:

$ sudo -u appuser jcmd $(pgrep -f orders.jar) VM.native_memory summary
1210:
Native memory tracking is not enabled

On a systemd service the least invasive way is a drop-in:

$ sudo systemctl edit orders
#   [Service]
#   Environment=JAVA_TOOL_OPTIONS=-XX:NativeMemoryTracking=summary
$ sudo systemctl restart orders
$ journalctl -u orders -n 20 | grep Picked
java[17390]: Picked up JAVA_TOOL_OPTIONS: -XX:NativeMemoryTracking=summary

Reading the summary

$ sudo -u appuser jcmd $(pgrep -f orders.jar) VM.native_memory summary
17390:

Native Memory Tracking:

(Omitting categories weighting less than 1KB)

Total: reserved=2127809KB, committed=598240KB
       malloc: 29095KB #165519
       mmap:   reserved=2098714KB, committed=569145KB

-                 Java Heap (reserved=524288KB, committed=370688KB)
                            (mmap: reserved=524288KB, committed=370688KB)

-                     Class (reserved=1050624KB, committed=14720KB)
                            (classes #16234)
                            (  instance classes #15321, array classes #913)
                            ...
-                 Metaspace (reserved=100352KB, committed=99602KB)

-                    Thread (reserved=116480KB, committed=17472KB)
                            (thread #56)
                            (stack: reserved=114240KB, committed=15680KB)

-                      Code (reserved=253952KB, committed=46080KB)

-                        GC (reserved=39649KB, committed=22536KB)
-                     Other (reserved=16384KB, committed=16384KB)
-                    Symbol (reserved=20480KB, committed=20480KB)
...

Baseline and diff: what is growing

One summary is a photograph. What you want in an incident is a diff:

$ sudo -u appuser jcmd $(pgrep -f orders.jar) VM.native_memory baseline
PID:
Baseline taken
  ... wait a few minutes of real traffic ...
$ sudo -u appuser jcmd $(pgrep -f orders.jar) VM.native_memory summary.diff
PID:

Native Memory Tracking:

Total: reserved=2489602KB +361793KB, committed=951402KB +210310KB

-                 Java Heap (reserved=1075200KB, committed=1075200KB)
-                    Thread (reserved=895232KB +340000KB, committed=213888KB +52080KB)
                            (thread #422 +160)
...

Every changed category carries a + or - delta. Here Thread grew by 160 threads and 52 MB committed in a few minutes while the heap did not move at all. That is a thread leak, and it kills the container by RSS while every heap dashboard looks fine.

Mapping symptoms to categories

Thread grows, thread # grows     thread leak: executors created per request, never shut down
Other grows                      direct buffers: Netty/NIO, missing release(), MaxDirectMemorySize unset
Class/Metaspace grows            classloader leak (redeploys, dynamic proxies, scripting engines)
Code grows                       very rarely an issue; code cache full logs a warning
GC grows with the heap           normal
Nothing grows but RSS does       native code outside NMT: JNI libraries, glibc arenas (MALLOC_ARENA_MAX)

The OOM kill you cannot see from the heap

When the cgroup kills a JVM, nothing is logged by the JVM - SIGKILL gives it no chance. No heap dump, no OutOfMemoryError. The evidence is in the kernel log and the unit:

# during a cgroup OOM loop, like chapter 4's payments incident:
journalctl -k | grep -i 'out of memory'
... Memory cgroup out of memory: Killed process 5123 (java) total-vm:3412044kB, anon-rss:1391208kB, ...
systemctl status payments | grep -E 'Active|Main PID'
     Active: activating (auto-restart) (Result: oom-kill) since ...

Contrast with a heap OOM: the JVM itself throws java.lang.OutOfMemoryError: Java heap space, logs it, writes a heap dump if asked, and exits only if ExitOnOutOfMemoryError is set. Two different failures:

heap OOM     JVM throws OutOfMemoryError     app log, hprof file       exit code 3 (with ExitOnOOM) or limps on
cgroup OOM   kernel sends SIGKILL            kernel log, memory.events  exit 137, Result: oom-kill

"OOMKilled but the heap is only 60%" is always the second kind, and NMT is how you find what grew.

Why it helps

The hardest Java incidents are the ones where the pod dies but every heap graph looks fine: exit 137, Result: oom-kill, no OutOfMemoryError, no heap dump. Developers look at their heap graphs, see nothing and blame Kubernetes. NMT is how you end that argument with numbers.

With a baseline and a diff you can say "Thread grew by 160 threads and 52 MB in five minutes, the heap did not move", which points straight at an executor created per request. Or "Other is growing", which means direct buffers in a Netty client. This is also how you size a JVM's limit on evidence rather than a rule of thumb: the summary tells you exactly how much non-heap a service needs at peak.

Commands in this lesson

jcmd systemctl journalctl

FAQ

Can I turn NMT on for a running JVM?

No. NMT has to be enabled at startup with -XX:NativeMemoryTracking=summary, because it wraps allocations from the first moment. Without it, jcmd PID VM.native_memory summary just replies Native memory tracking is not enabled. The least invasive way on a systemd service is a drop-in with Environment=JAVA_TOOL_OPTIONS=-XX:NativeMemoryTracking=summary and a restart; in Kubernetes, the same env var in the pod spec and a rollout.

How much overhead does NMT add?

In summary mode, usually 1-2% CPU and a little memory for its own bookkeeping; worst case around 5-10% for allocation-heavy native workloads. Detail mode records a call site per allocation and costs more; it is meant for JVM developers and rarely needed for operations. Many teams leave summary on permanently for important services, so the data is there when an incident starts.

Why doesn't NMT's total match RSS exactly?

Two reasons. Committed memory is not always resident: the JVM can commit pages that were never touched, or that the OS swapped out, so committed can be higher than RSS. And NMT only sees the JVM's own allocations: glibc's malloc arenas, JNI libraries and native code linked into the process are outside its view. A small gap is normal. A gap that grows over time points to native code or malloc arenas.

What does "reserved" mean compared with "committed"?

Reserved is address space the JVM has claimed but not necessarily backed with memory, like the 1 GB for the compressed class space or 240 MB for the code cache. It costs nothing and explains the large VSZ. Committed is memory the JVM has actually asked the OS to back, and it is the column you compare with RSS and the container limit. Ignore reserved when sizing.

Which category grows in which kind of leak?

Thread with a rising thread # is a thread leak, usually an executor created per request and never shut down. Other growing is direct buffers, from Netty, NIO or a missing release(). Class or Metaspace with a rising classes # is a classloader leak or runtime class generation. GC grows with the heap and is normal. If nothing grows but RSS does, the leak is outside NMT: JNI or glibc arenas.

In an interview Mid

What is Native Memory Tracking and when would you use it?

NMT is the JVM's own accounting of its native memory, per category: Java Heap, Class/Metaspace, Thread, Code, GC, Other... I use it when RSS grows or the pod is OOMKilled while the heap looks fine.

Reading the growth: Thread (thread # climbing) = a thread leak; Other = direct buffers; Class/Metaspace = a classloader leak; nothing grows but RSS does = native code outside NMT (JNI, glibc arenas).

Also asked: What memory areas does a JVM have besides the heap? · How do you tell a heap OutOfMemoryError from a cgroup OOM kill? · Why does a Java pod get OOMKilled with no OutOfMemoryError in its logs?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.