JVM Internals: interview questions
The question you are most likely to get for each topic, a model answer, and what else comes up. From chapter 20 of the course.
A Java pod keeps getting OOMKilled but the heap graphs look fine. How do you investigate? Mid
Exit 137 with a healthy heap is a cgroup OOM caused by memory outside the heap, so I look at the whole process.
- Confirm the kill.
kubectl describe podshowsOOMKilled, exit 137; on a VMjournalctl -k | grep -i 'out of memory'andResult: oom-kill. The JVM logs nothing - SIGKILL gives it no chance, so noOutOfMemoryErrorand no heap dump. - Compare the limit with the heap.
jcmd 1 VM.flagsshowsMaxHeapSize.-Xmxat or near the limit is the classic cause: metaspace, thread stacks, code cache and direct buffers add 100-400 MB. - If the sizing looks right, measure. Start with
-XX:NativeMemoryTracking=summary(viaJAVA_TOOL_OPTIONS), take aVM.native_memory baseline, later asummary.diff: Thread growing = thread leak, Other = direct buffers, Class/Metaspace = classloader leak. - Fix:
-XX:MaxRAMPercentage=75instead of-Xmx, a limit of about heap x 1.3-1.5, and the leak NMT pointed at.
Also asked: What is the difference between a heap OOM and an OOM kill? · How would you size the memory limit for a JVM running in a container? · A Java service is at 100% CPU with normal traffic. What do you check first?
Why does a Java process use more memory than its -Xmx? Mid
Because -Xmx limits only the Java heap, and the JVM is a native program with other allocations that the kernel, the cgroup and the OOM killer all count:
- Metaspace - class metadata, not limited by
-Xmx(a Spring app with 16 000 classes needs ~100 MB). - Thread stacks - one per thread (
NLWPinps). - Code cache - the JIT's compiled code.
- GC structures, symbols, and direct buffers (which default to about the heap size).
On oncall-lab orders runs with -Xmx512m and ps -o pid,rss,vsz,nlwp shows ~597 MB RSS, while jcmd 1210 GC.heap_info shows 362 MB committed heap. 100-400 MB of non-heap is normal; a leak is RSS or after-GC heap that keeps climbing.
For operations: the container limit must hold the whole process - rule of thumb limit = heap x 1.3 to 1.5, then measure. VSZ (3.4 GB here) is reserved address space and costs nothing.
Also asked: What is the difference between RSS and VSZ for a Java process? · What are the young and old generations, and why does the heap have them? · Heap usage is at 90%. Should you be worried?
Learn it: 20.1 The heap is not the footprint
How does the JVM decide its heap size in a container, and what would you set? Mid
With container support (-XX:+UseContainerSupport, on by default since JDK 10) the JVM reads the memory and CPU limits from the cgroup and computes its defaults - ergonomics:
- Heap =
MaxRAMPercentage, default 25% of the limit: a 1 GB container gets a 256 MB heap. - Below 2 CPUs or ~1792 MB it is not a "server-class machine" and silently picks the Serial collector instead of G1.
What I set: -XX:MaxRAMPercentage=75 (the heap follows the limit if someone changes it, unlike -Xmx), -XX:+UseG1GC explicitly, and never a heap equal to the limit - the rest of the footprint needs ~25%. Below ~512 MB, 50-60%.
How I verify: java -XX:+PrintFlagsFinal -version (the {ergonomic} / {command line} column says where each value came from), -XshowSettings:system (Memory Limit: Unlimited in a container = an old JDK sizing from the node), and Picked up JAVA_TOOL_OPTIONS in the log.
Also asked: Why prefer MaxRAMPercentage over -Xmx in a container? · How do you pass JVM flags to a service without changing its start command? · Why would a JVM in a small container end up on the Serial collector?
Learn it: 20.2 What the JVM decides at startup: container-aware ergonomics
How do you take a thread dump of a Java process, and why might it fail? Mid
sudo -u appuser jcmd PID Thread.print > dump1.txt- the preferred way (jstack -l PIDis the older tool, same content).kill -3 PID(SIGQUIT) - the JVM prints the dump to its own stdout: the journal for a systemd service,kubectl logsin a pod. Needs no JDK tools.
jcmd, jstack and jmap use the attach mechanism (a SIGQUIT plus a UNIX socket in the JVM's /tmp), which is why they fail:
- Wrong user -
Operation not permitted. Run as the JVM's user (sudo -u appuser) or root. - Different /tmp -
PrivateTmp=yesor a container; the tool looks under/proc/PID/root/tmp. - JVM not answering (hung in GC, out of memory) -
AttachNotSupportedException. - No tools at all - the image has only the JRE:
jcmdnot found. Use a JDK image or a debug container withkubectl debug --target.
Also asked: What is the difference between the JRE and the JDK, and why does it matter on servers? · Where does a heap dump taken with jcmd get written? · Why does jstat say "not found" for a process that is running?
Learn it: 20.3 The JDK tools, and who may attach
What is Native Memory Tracking and when would you use it? Mid
NMT is the JVM's own accounting of its native memory, per category: Java Heap, Class/Metaspace, Thread, Code, GC, Other... I use it when RSS grows or the pod is OOMKilled while the heap looks fine.
- It must be on at startup:
-XX:NativeMemoryTracking=summary, for a service via a drop-inEnvironment=JAVA_TOOL_OPTIONS=...and a restart (Picked up JAVA_TOOL_OPTIONSin the journal proves it). Overhead is usually ~1-2%. jcmd PID VM.native_memory summary- compare Total committed with RSS; reserved is just address space.- The useful part is the diff:
VM.native_memory baseline, wait under real traffic, thensummary.diff.
Reading the growth: Thread (thread # climbing) = a thread leak; Other = direct buffers; Class/Metaspace = a classloader leak; nothing grows but RSS does = native code outside NMT (JNI, glibc arenas).
Also asked: What memory areas does a JVM have besides the heap? · How do you tell a heap OutOfMemoryError from a cgroup OOM kill? · Why does a Java pod get OOMKilled with no OutOfMemoryError in its logs?
Learn it: 20.8 Native memory tracking: where the rest of the RSS goes
What is the difference between a minor GC and a full GC, and when is GC a problem? Mid
Most objects die young, so the heap has a young generation (eden + survivors) and an old generation.
- Young (minor) GC - copies the few live objects out of eden and throws eden away. Cost is proportional to what survives, so it is cheap (2-20 ms) and frequent. A young GC every few seconds under load is healthy.
- Full GC - the whole heap, compacting, stop-the-world: every application thread frozen at a safepoint, from 100 ms to seconds. With G1 (the default, pause target 200 ms) frequent Full GCs are a symptom: the live set fills the heap - a leak, a heap too small, or a burst that promoted too much.
GC is a problem when it shows up from outside: latency spikes that line up with pauses (a 400 ms pause adds 400 ms to every in-flight request), or 100% CPU with throughput near zero - Full GC after Full GC freeing almost nothing. A low-pause collector like ZGC does not fix a leak.
Also asked: What is a stop-the-world pause? · Compare the Serial, Parallel, G1 and ZGC collectors at a high level. · Which GC-related JVM flags would you put on every production service?
Learn it: 20.12 Garbage collection: pauses, minor vs full, and the collectors
How would you tell from a GC log whether a service has a memory leak? Mid
Turn the log on with unified logging, rotated: -Xlog:gc*:file=/var/log/app/gc.log:time,uptime:filecount=5,filesize=10M.
Each collection ends in a summary line like 223M->165M(272M) 4.900ms: heap used before -> after (committed), pause.
Used goes up and down all day, so ignore it. Look at the heap after collections over hours - ideally after Mixed or Full GCs, which include the old generation:
- After-GC floor keeps climbing (100 -> 234 -> 310 -> 408 MB over five hours, committed growing to the maximum) = a leak: something keeps objects reachable.
- After-GC floor moves in a band and returns to a baseline = churn, healthy.
- Caveat: a cache filling to its configured size climbs too, then flattens - look at a longer window.
The end state is lines like Pause Full ... 511M->498M(512M) 431ms: a long stop-the-world pause that freed almost nothing, right before OutOfMemoryError.
Also asked: How do you enable GC logging on a JVM in production? · Latency spikes every few minutes on a Java service. How do you check whether GC is responsible? · What does a Full GC that frees almost nothing tell you?
A Java application is at 100% CPU. What steps do you take? Mid
- Is it the GC?
sudo -u appuser jstat -gcutil PID 1s 10. The counters are cumulative, so watch the rates:FGC/FGCTclimbing every second andO(old gen) stuck near 100% = the collector is thrashing. Then it is a heap too small for the live set, or a leak - go to the histogram. - Is it one thread?
top -H -p PIDshows threads by CPU; find that thread id asnid=in a thread dump (printf '%x\n' 1297on older JDKs that print hex). - Is it all the request threads? Many
RUNNABLEthreads in your own code: real load or an inefficient loop. A profiler says where:jcmd PID JFR.start duration=60s filename=/tmp/cpu.jfr, or async-profiler. - Is it the JIT?
C2 CompilerThread0busy right after startup is warm-up and settles.
jstat reads hsperfdata without attaching, so the wrong user gets PID not found.
Also asked: What do the main columns of jstat -gcutil tell you? · How do you find which Java thread is using the CPU? · What can jstat not tell you about a Java process?
Learn it: 20.17 jstat, and "is it our code or the collector?"
What does a Java thread dump contain, and what do you look at first? Mid
A snapshot of every thread: for each one the name, Java id, nid (the OS thread id top -H shows), cpu= used so far, the java.lang.Thread.State, the stack (most recent frame first) and lock lines (waiting to lock <0x...>, locked <0x...>). It pauses the JVM for milliseconds - safe in production.
What I look at first:
- Names -
http-nio-8080-exec-*are Tomcat request threads,HikariPool-1the DB pool,pool-N-thread-Man unnamed executor. - Counts, not individuals -
grep 'java.lang.Thread.State' dump1.txt | sort | uniq -c | sort -rn, then state + top frame. Fifty threads with the same stack is the answer. - The RUNNABLE trap - a thread waiting for bytes in
Net.pollis RUNNABLE, not busy;cpu=across two dumps settles it. - Same lock address in many threads - they all wait for one thing; search for who
lockedit.
Also asked: Explain the Java thread states and what each one suggests during an incident. · Is taking a thread dump safe on a production service? · How do you match a thread that is hot in top -H to a thread in a dump?
Requests to a Java service hang, CPU is low and there are no errors. How do you find out why? Mid
Take three thread dumps ten seconds apart (jcmd PID Thread.print > dump$i.txt) and count clusters of identical stacks. Same thread, same stack, same cpu= in all three = stuck. Then match the pattern:
- Connection pool exhaustion - 180 request threads
TIMED_WAITINGinHikariPool.getConnection, and a small cluster (20 threads = the pool size) holding connections while waiting on a downstream call. The big cluster is the symptom, the small one the cause. Presents as latency, not errors. - Downstream hang - all request threads
RUNNABLEinNet.pollinside the same client class: a blocking read with no read timeout (read.timeout.ms=0) waits forever. - Deadlock - the dump ends with
Found one Java-level deadlock; both threadsBLOCKED (on object monitor). - Async pool exhaustion - Tomcat idle, every executor thread stuck in the same call.
Fix is usually a timeout; for a deadlock, keep the dump, restart, hand the stacks to the developers.
Also asked: What is a deadlock and how do you detect one in Java? · Why take three thread dumps instead of one? · How can one slow dependency take down a whole service?
Learn it: 20.22 Thread dump patterns: clusters, exhaustion, deadlock, async
How would you find the cause of a memory leak in a Java service? Mid
- Confirm it - the heap after GC keeps climbing (GC log, or
jstat'sOfloor). - What is accumulating - two histograms a few minutes apart under traffic:
jmap -histo:live PID > h1.txt, again intoh2.txt. The leak is the class whose count grows every time and never falls - look for your own packages (lab.orders.cache.OrderSnapshot48k -> 91k), not[BandString, which top every heap. - Who is holding it - a heap dump,
jcmd PID GC.heap_dump /tmp/orders.hprof, opened in Eclipse MAT: Leak Suspects, the dominator tree by retained size, and the path to GC roots (static field -> map -> entry) - that chain is the bug report.
Before dumping: it pauses the JVM, is as big as the live heap (pick a disk with room), and contains customer data and secrets (mode 0600; do not attach it to a ticket). Set -XX:+HeapDumpOnOutOfMemoryError so the next failure dumps itself.
Also asked: How can Java leak memory when it has a garbage collector? · What precautions do you take before taking a heap dump in production? · What is the difference between shallow size and retained size?
Learn it: 20.28 Heap dumps and finding a leak
How do you troubleshoot a Java application running in a Kubernetes pod? Mid
The same tools, with kubectl exec POD -- in front and PID 1 (the JVM is the entrypoint): kubectl exec POD -- jcmd 1 Thread.print > dump1.txt, jcmd 1 GC.heap_info, jcmd 1 VM.flags. It works because exec runs as the JVM's user, in the same /tmp.
When the image has no tools (JRE or distroless, executable file not found):
kubectl exec POD -- kill -3 1and readkubectl logs(needs akillbinary).- An ephemeral container:
kubectl debug -it POD --image=eclipse-temurin:21-jdk --target=app- it shares the process namespace; it must run as the same UID or attach fails.
Files out: kubectl cp NAMESPACE/POD:/tmp/heap.hprof . - write dumps to a real volume, not a memory-backed emptyDir.
And check the flags: JAVA_TOOL_OPTIONS with MaxRAMPercentage=75, ExitOnOutOfMemoryError (or a half-dead JVM keeps passing liveness), and CPU limits: limits.cpu: 1 means Serial GC and a one-thread common pool.
Also asked: What happens to a JVM when its pod exceeds the memory limit? · Why does ExitOnOutOfMemoryError matter more in Kubernetes than on a VM? · How do CPU limits change what the JVM decides at startup?
Learn it: 20.35 The same tools in Kubernetes, and the images you cannot debug
Practise these answers with flashcards and labs Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.