OnCallReady

Lesson 20.28 · JVM Internals · 14 min read

Heap dumps and finding a leak

In plain words

Imagine a class where every pupil must hand in their homework, and the teacher puts it on a shelf "just in case". Nobody ever takes anything off. After a term the shelf is overflowing. The cleaner cannot throw any of it away, because it is all still on the teacher's shelf, so technically it is still wanted.

A Java memory leak is exactly that: objects that stay reachable, usually in a collection that only grows. The garbage collector cannot remove them. A histogram (jmap -histo:live) is counting what is on the shelf by type; comparing two counts shows OrderSnapshot going from 48 000 to 91 000. A heap dump opened in Eclipse MAT shows whose shelf it is: a static OrderSnapshotCache with no eviction.

Is it a leak?

The problem. The heap after each GC climbs for hours until the service dies with OutOfMemoryError. Something keeps objects alive that should have been freed - a memory leak. You find it by counting objects by class and following who holds them.

What you need to know already: the after-GC trend (20.13), jcmd GC.class_histogram and GC.heap_dump (20.3), sort/diff (7.4).

A heap dump is a file containing every object in the heap and who points to whom; a histogram is the much smaller table "class, number of objects, bytes".

Established by the GC log (after-GC floor climbing) or jstat (O's floor climbing). A leak is objects that stay reachable - usually a collection that only ever grows: a static map used as a cache, a listener list, a ThreadLocal never cleared, a queue nobody drains.

Histograms: the cheap first look

$ sudo -u appuser jmap -histo:live $(pgrep -f orders.jar) | head -12
 num     #instances         #bytes  class name (module)
-------------------------------------------------------
   1:        304046      132152152  [B ([email protected])
   2:        196812        4723488  java.lang.String ([email protected])
   3:         91601        2931232  java.util.HashMap$Node ([email protected])
   4:         43613        2704008  [Ljava.lang.Object; ([email protected])
   5:         60841        1946912  java.util.concurrent.ConcurrentHashMap$Node ([email protected])
   6:         17038        1908256  java.lang.Class ([email protected])
   ...
Total       1832711      178345678
num          rank by bytes
#instances   live objects of that class
#bytes       SHALLOW size: the objects themselves, not what they point to
class name   [B = byte[], [I = int[], [Ljava.lang.Object; = Object[], (module) for JDK classes

The technique is the diff: two histograms, a few minutes apart, under traffic. The leak is the class whose count grows every time and never falls:

$ sudo -u appuser jmap -histo:live $(pgrep -f orders.jar) > h1.txt
$ sleep 120
$ sudo -u appuser jmap -histo:live $(pgrep -f orders.jar) > h2.txt
$ grep -E 'lab\.' h1.txt h2.txt
h1.txt:  14:         48412        2323776  lab.orders.cache.OrderSnapshot
h2.txt:   9:         91530        4393440  lab.orders.cache.OrderSnapshot

An application class climbing from 48 000 to 91 000 instances in two minutes, while traffic is steady. Look for your own packages (lab.orders here) and for the collections that hold them (ConcurrentHashMap$Node growing in step). A per-class diff with join:

join -1 4 -2 4 <(awk 'NR>2{print $1,$2,$3,$4}' h1.txt | sort -k4) \
               <(awk 'NR>2{print $1,$2,$3,$4}' h2.txt | sort -k4) \
  | awk '{d=$6-$3; if (d>1000) print d, $1}' | sort -rn | head

Heap dumps: the full picture

A histogram says what is accumulating. A heap dump says who is holding it.

$ sudo -u appuser jcmd $(pgrep -f orders.jar) GC.heap_dump /tmp/orders.hprof
1210:
Dumping heap to /tmp/orders.hprof ...
Heap dump file created [301451712 bytes in 0.793 secs]

$ ls -l /tmp/orders.hprof
-rw------- 1 appuser appuser 301451712 Sep 23 10:22 /tmp/orders.hprof

Know before you take one:

Automatically, at the moment it matters:

-XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=/var/tmp/
java.lang.OutOfMemoryError: Java heap space
Dumping heap to /var/tmp/java_pid1210.hprof ...
Heap dump file created [538012443 bytes in 1.402 secs]

HeapDumpOnOutOfMemoryError is manageable: you can switch it on in a running JVM without a restart - jinfo -flag +HeapDumpOnOutOfMemoryError PID or jcmd PID VM.set_flag HeapDumpOnOutOfMemoryError true.

Analysing it: Eclipse MAT

Copy the file to a machine with memory to spare (scp, or kubectl cp ns/pod:/tmp/orders.hprof ./orders.hprof) and open it in Eclipse Memory Analyzer (MAT). The Leak Suspects report does most of the work:

Problem Suspect 1
One instance of "lab.orders.cache.OrderSnapshotCache" loaded by "jdk.internal.loader.ClassLoaders$AppClassLoader"
occupies 246,512,112 (81.72%) bytes. The memory is accumulated in one instance of
"java.util.concurrent.ConcurrentHashMap$Node[]".
Keywords: lab.orders.cache.OrderSnapshotCache, java.util.concurrent.ConcurrentHashMap$Node[]

The concepts MAT uses:

What you tell the developers

Not "the heap is full". This:

orders leaks ~25 MB/min under normal traffic (GC log: after-GC floor 150 -> 420 MB in 11 min)
jmap -histo:live: lab.orders.cache.OrderSnapshot 48k -> 91k instances in 2 min, never falls
MAT: 82% of the heap retained by static OrderSnapshotCache.CACHE (ConcurrentHashMap, no eviction)
heap dump: /var/tmp/orders-2026-09-23.hprof (restricted, contains customer data)
mitigation: restart every 6h until fixed; HeapDumpOnOutOfMemoryError + ExitOnOutOfMemoryError now set

Why it helps

When you have established that a service leaks, the next question from everyone is "what?". A histogram diff answers it in two minutes without heavy tools, and a heap dump plus MAT's Leak Suspects report points at the exact field that holds the memory. You give developers a class name, a growth rate and a reference chain instead of "the heap is full".

It also matters that you know the costs: a live histogram forces a full GC, and a heap dump pauses the JVM and writes a file as big as the heap that contains passwords and customer data. At a bank, handling that file is a security question, and you will be the one who knows it must not be attached to a ticket.

Commands in this lesson

jmap sleep grep jcmd ls

FAQ

Why are byte[] and String always at the top of the histogram?

Because almost everything in Java contains strings, and strings are backed by byte arrays. Every JSON payload, header, SQL string and log message ends up there. Their presence at the top says nothing on its own. The technique is the diff: take two histograms a few minutes apart under steady traffic and look for classes, especially from your own packages like lab.orders, whose counts grow every time and never fall.

What is the difference between shallow and retained size?

Shallow size is the memory of the object itself, which is what the histogram shows. Retained size is everything that would be freed if that object were collected, including all the objects only it references. A HashMap object is tiny in shallow terms but can retain 250 MB of entries. Leaks are found by retained size, which is why MAT's dominator tree ranks objects that way.

What does the :live option do to the application?

jmap -histo:live and jcmd GC.class_histogram force a full GC before counting, so the numbers only include reachable objects. That full GC is a stop-the-world pause: fine on a 500 MB heap, noticeable on a 30 GB one. Without :live, garbage is counted too and the numbers jump around from one run to the next, which makes diffs useless. Choose live, but pick your moment on large heaps.

How big is a heap dump and where should I write it?

About as big as the live heap: a 12 GB heap makes a file of roughly 12 GB. Write it somewhere with room, as an absolute path the JVM's user can write. In a container /tmp is often the small writable layer, and a memory-backed emptyDir counts against the memory limit, so a dump there can get the pod OOMKilled while dumping. Use a real volume and copy it out with kubectl cp or scp.

Is a heap dump sensitive?

Yes, extremely. It contains everything in the JVM's memory: passwords, API tokens, session data, customer names and account numbers. Treat it like a database backup. That is why the JVM creates it with mode 0600. Do not attach it to a ticket, do not copy it to your laptop without approval, and delete it when the analysis is done. Share the MAT report and the reference chain instead.

In an interview Mid

How would you find the cause of a memory leak in a Java service?

  1. Confirm it - the heap after GC keeps climbing (GC log, or jstat's O floor).
  2. What is accumulating - two histograms a few minutes apart under traffic: jmap -histo:live PID > h1.txt, again into h2.txt. The leak is the class whose count grows every time and never falls - look for your own packages (lab.orders.cache.OrderSnapshot 48k -> 91k), not [B and String, which top every heap.
  3. Who is holding it - a heap dump, jcmd PID GC.heap_dump /tmp/orders.hprof, opened in Eclipse MAT: Leak Suspects, the dominator tree by retained size, and the path to GC roots (static field -> map -> entry) - that chain is the bug report.

Before dumping: it pauses the JVM, is as big as the live heap (pick a disk with room), and contains customer data and secrets (mode 0600; do not attach it to a ticket). Set -XX:+HeapDumpOnOutOfMemoryError so the next failure dumps itself.

Also asked: How can Java leak memory when it has a garbage collector? · What precautions do you take before taking a heap dump in production? · What is the difference between shallow size and retained size?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.