Chapter 20 JVM Internals
Memory layout and container sizing, the JDK tools, GC logs and jstat, native memory, thread dumps, deadlocks and leaks - on the Java services already running on this box.
In plain words
Think of a busy school canteen. The kitchen staff cook, but there is also a cleaner who walks around picking up plates people have finished with. If the cleaner is slow, tables fill up with dirty plates and nobody new can sit down. If the cleaner has to stop everyone eating so he can clean properly, the whole room freezes for a moment. And the canteen needs space for more than tables: a store room, a coat rack, a kitchen.
A Java service is that canteen. The heap is the tables, the garbage collector is the cleaner, pauses are the frozen moments, and the store room and kitchen are metaspace, thread stacks and code cache. This chapter teaches you to look inside orders (PID 1210) and payments with jcmd, jstat, GC logs, thread dumps and heap dumps.
Why it matters on call
Most of the services you will run at a bank are Spring Boot on a JVM, and most of their incidents look like platform problems while being JVM problems: a pod OOMKilled with exit 137 while the heap graph sits at 60%, a service at 100% CPU that is really the collector thrashing, p99 going vertical while CPU is idle because every thread is waiting on a pool. Developers will ask the platform team "is it Kubernetes?" and you need to answer with evidence.
This chapter comes after memory, cgroups, containers and Kubernetes because it builds on all of them: the cgroup limit from Ch 5, the container image from Ch 10, the pod spec from Ch 15. Next comes Spring Boot's runtime side, which reads the same problems from the outside through Actuator and metrics.
Lessons
- The heap is not the footprint
- What the JVM decides at startup: container-aware ergonomics
- The JDK tools, and who may attach
- Native memory tracking: where the rest of the RSS goes
- Garbage collection: pauses, minor vs full, and the collectors
- Reading unified GC logs, and the leak test
- jstat, and "is it our code or the collector?"
- Reading a thread dump: format and states
- Thread dump patterns: clusters, exhaustion, deadlock, async
- Heap dumps and finding a leak
- The same tools in Kubernetes, and the images you cannot debug
24 hands-on labs (missions, incidents and drills) run in the terminal: Open this chapter in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.
Questions people ask
Do I need to know Java to do this chapter?
No. You need to read what the JVM tells you, not write Java. A thread dump is a list of stacks with names and states; a GC log is one line per collection with before, after and pause time; a histogram is a table of classes by count. Reading package names like lab.orders.cache helps you point developers at their own code, but the diagnosis itself is pattern matching on tool output, the same skill as reading ps or journalctl.
Why is the process bigger than -Xmx?
Because -Xmx only limits the Java heap, and the JVM is a native program with many other allocations: metaspace for class metadata, a stack per thread, the JIT's code cache, GC bookkeeping, direct buffers and malloc arenas. On this box orders runs with -Xmx512m and has about 590 MB resident. A gap of 100-400 MB is normal. What matters is whether it keeps climbing, and NMT tells you which category does.
Which tool do I use for what?
jcmd does almost everything: flags, heap info, thread dumps, heap dumps, NMT and JFR. jstat is the live counter view of the collector, good for "is it GC right now?". GC logs give you history, and the after-GC floor is the leak test. jmap -histo:live shows which classes are accumulating. Eclipse MAT opens a heap dump and tells you who is holding the memory. Actuator gives you dumps over HTTP when the image has no JDK.
Why does jcmd say "Operation not permitted" even with the right PID?
The attach mechanism sends the JVM a signal and talks over a UNIX socket in the JVM's /tmp. You must be the same user as the JVM, or root, and see the same /tmp. orders runs as appuser, so sudo -u appuser jcmd 1210 ... works and plain jcmd from learner does not. jstat fails differently, with 1210 not found, because it reads the per-user hsperfdata file instead of attaching.
Is the JVM on my Mac the same as in the container?
The JVM is the same HotSpot, but what it decides at startup differs. On your Mac it sees all your RAM and cores; in a container it reads the cgroup limits and sizes the heap at 25% of the limit by default, and below 2 CPUs or about 1.8 GB it silently picks the Serial collector. The aarch64 default thread stack is also 2 MB instead of x86's 1 MB. Always check with -XX:+PrintFlagsFinal or jcmd PID VM.flags where it runs.