Is it a leak?
The problem. The heap after each GC climbs for hours until the service dies with OutOfMemoryError. Something keeps objects alive that should have been freed - a memory leak. You find it by counting objects by class and following who holds them.
What you need to know already: the after-GC trend (20.13), jcmd GC.class_histogram and GC.heap_dump (20.3), sort/diff (7.4).
A heap dump is a file containing every object in the heap and who points to whom; a histogram is the much smaller table "class, number of objects, bytes".
Established by the GC log (after-GC floor climbing) or jstat (O's floor climbing). A leak is objects that stay reachable - usually a collection that only ever grows: a static map used as a cache, a listener list, a ThreadLocal never cleared, a queue nobody drains.
Histograms: the cheap first look
$ sudo -u appuser jmap -histo:live $(pgrep -f orders.jar) | head -12
num #instances #bytes class name (module)
-------------------------------------------------------
1: 304046 132152152 [B ([email protected])
2: 196812 4723488 java.lang.String ([email protected])
3: 91601 2931232 java.util.HashMap$Node ([email protected])
4: 43613 2704008 [Ljava.lang.Object; ([email protected])
5: 60841 1946912 java.util.concurrent.ConcurrentHashMap$Node ([email protected])
6: 17038 1908256 java.lang.Class ([email protected])
...
Total 1832711 178345678
num rank by bytes
#instances live objects of that class
#bytes SHALLOW size: the objects themselves, not what they point to
class name [B = byte[], [I = int[], [Ljava.lang.Object; = Object[], (module) for JDK classes
:liveforces a Full GC first and counts only reachable objects. It pauses the application for the length of that GC - fine on a 500 MB heap, think twice on a 30 GB one. Without:livegarbage is counted too and numbers jump around.jcmd PID GC.class_histogramis the same (live by default,-allfor everything).[BandStringare at the top of every Java heap. That alone means nothing. Strings are backed by byte arrays; everything contains strings.
The technique is the diff: two histograms, a few minutes apart, under traffic. The leak is the class whose count grows every time and never falls:
$ sudo -u appuser jmap -histo:live $(pgrep -f orders.jar) > h1.txt
$ sleep 120
$ sudo -u appuser jmap -histo:live $(pgrep -f orders.jar) > h2.txt
$ grep -E 'lab\.' h1.txt h2.txt
h1.txt: 14: 48412 2323776 lab.orders.cache.OrderSnapshot
h2.txt: 9: 91530 4393440 lab.orders.cache.OrderSnapshot
An application class climbing from 48 000 to 91 000 instances in two minutes, while traffic is steady. Look for your own packages (lab.orders here) and for the collections that hold them (ConcurrentHashMap$Node growing in step). A per-class diff with join:
join -1 4 -2 4 <(awk 'NR>2{print $1,$2,$3,$4}' h1.txt | sort -k4) \
<(awk 'NR>2{print $1,$2,$3,$4}' h2.txt | sort -k4) \
| awk '{d=$6-$3; if (d>1000) print d, $1}' | sort -rn | head
Heap dumps: the full picture
A histogram says what is accumulating. A heap dump says who is holding it.
$ sudo -u appuser jcmd $(pgrep -f orders.jar) GC.heap_dump /tmp/orders.hprof
1210:
Dumping heap to /tmp/orders.hprof ...
Heap dump file created [301451712 bytes in 0.793 secs]
$ ls -l /tmp/orders.hprof
-rw------- 1 appuser appuser 301451712 Sep 23 10:22 /tmp/orders.hprof
Know before you take one:
- It pauses the JVM for the dump's duration (seconds per GB) and triggers a Full GC first (use
-allto skip it and include garbage). - It is as big as the live heap. A 12 GB heap makes a ~12 GB file - do not write it to a 10 GB container overlay.
- It contains everything in memory: passwords, tokens, customer data. Treat it like a database backup. Mode 0600 for a reason; do not attach it to a ticket.
- The target JVM writes it, as its user, relative to its working directory.
Automatically, at the moment it matters:
-XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=/var/tmp/
java.lang.OutOfMemoryError: Java heap space
Dumping heap to /var/tmp/java_pid1210.hprof ...
Heap dump file created [538012443 bytes in 1.402 secs]
HeapDumpOnOutOfMemoryError is manageable: you can switch it on in a running JVM without a restart - jinfo -flag +HeapDumpOnOutOfMemoryError PID or jcmd PID VM.set_flag HeapDumpOnOutOfMemoryError true.
Analysing it: Eclipse MAT
Copy the file to a machine with memory to spare (scp, or kubectl cp ns/pod:/tmp/orders.hprof ./orders.hprof) and open it in Eclipse Memory Analyzer (MAT). The Leak Suspects report does most of the work:
Problem Suspect 1
One instance of "lab.orders.cache.OrderSnapshotCache" loaded by "jdk.internal.loader.ClassLoaders$AppClassLoader"
occupies 246,512,112 (81.72%) bytes. The memory is accumulated in one instance of
"java.util.concurrent.ConcurrentHashMap$Node[]".
Keywords: lab.orders.cache.OrderSnapshotCache, java.util.concurrent.ConcurrentHashMap$Node[]
The concepts MAT uses:
- Shallow size - the object itself (what the histogram shows).
- Retained size - everything that would be freed if this object were collected. A small
HashMapobject can retain 250 MB. - Dominator tree - objects ranked by retained size. The leak is usually one entry at the top.
- Path to GC roots - the reference chain that keeps an object alive (
static field -> map -> entry -> your object). That chain is the bug report.
What you tell the developers
Not "the heap is full". This:
orders leaks ~25 MB/min under normal traffic (GC log: after-GC floor 150 -> 420 MB in 11 min)
jmap -histo:live: lab.orders.cache.OrderSnapshot 48k -> 91k instances in 2 min, never falls
MAT: 82% of the heap retained by static OrderSnapshotCache.CACHE (ConcurrentHashMap, no eviction)
heap dump: /var/tmp/orders-2026-09-23.hprof (restricted, contains customer data)
mitigation: restart every 6h until fixed; HeapDumpOnOutOfMemoryError + ExitOnOutOfMemoryError now set