OnCallReady

Lesson 20.3 · JVM Internals · 19 min read

The JDK tools, and who may attach

In plain words

Imagine a toy car with a little door on its belly for checking the batteries. To open it you need the right screwdriver, and it only fits if you are the owner of the car. If the car came from a shop that glued the door shut to make it look sleek, you cannot look inside at all.

The JDK tools, jcmd, jstack, jmap, jstat, are the screwdriver. They come with the JDK package, not the runtime (JRE), and they attach through a socket in the JVM's /tmp, only for the JVM's own user or root. That is why sudo -u appuser jcmd 1210 VM.flags works on oncall-lab, and why a distroless image with no JDK is the glued-shut car.

JRE vs JDK

The problem. A Java service misbehaves and ps, top and strace only show you "a process called java". The JDK ships tools that ask the JVM itself what is going on inside - if they are installed, and if you run them as the right user.

What you need to know already: users and sudo -u (4.3), signals and SIGQUIT (3.6), UNIX sockets and /proc/PID/root (3.14), systemd's PrivateTmp (2.26), JDK vs JRE (20.1).

The runtime (JRE) runs programs. The diagnostic tools - jcmd, jstack, jmap, jstat, jinfo, jps - ship with the JDK. On Ubuntu:

openjdk-21-jre-headless    java only                         what servers usually get
openjdk-21-jdk-headless    java + jcmd jstack jmap jstat ... what you want when debugging
$ jcmd
Command 'jcmd' not found, but can be installed with:
sudo apt install openjdk-21-jdk-headless
$ sudo apt install -y openjdk-21-jdk-headless

The same problem, worse, in containers: eclipse-temurin:21-jre and every distroless Java image have no jcmd at all. You have shipped an image you cannot look inside. Decide before 3am which of these you use:

  1. A JDK-based runtime image - bigger, but kubectl exec -it pod -- jcmd 1 ... works.
  2. A debug sidecar or ephemeral container with a JDK, sharing the process namespace: kubectl debug -it pod --image=eclipse-temurin:21-jdk --target=app. The target's PID 1 is visible from the debug container; attach needs the same UID and a shared /tmp, which is why this sometimes needs extra care.
  3. Actuator over HTTP (next chapter): /actuator/threaddump and /actuator/heapdump need no tools in the image at all.

How attaching works - and why it fails

jcmd, jstack, jmap and jinfo use the attach mechanism: the tool creates .attach_pid<PID>, sends the JVM SIGQUIT, the JVM opens a UNIX socket /tmp/.java_pid<PID> and the tool sends it a command. That imposes rules:

$ jcmd $(pgrep -f orders.jar) VM.flags
1210:
java.io.IOException: Operation not permitted
	at jdk.attach/sun.tools.attach.VirtualMachineImpl.sendQuitTo(Native Method)
	at jdk.attach/sun.tools.attach.VirtualMachineImpl.checkCatchesAndSendQuitTo(VirtualMachineImpl.java:398)
	...

The fix is to become the JVM's user: sudo -u appuser jcmd 1210 VM.flags. Root also works.

com.sun.tools.attach.AttachNotSupportedException: Unable to open socket file /proc/1210/root/tmp/.java_pid1210: target process 1210 doesn't respond within 10500ms or HotSpot VM not loaded

jstat is different: it reads the JVM's shared performance counters in /tmp/hsperfdata_<user>/<pid> (mode 0600) without attaching. Another user just gets:

$ jstat -gcutil $(pgrep -f orders.jar)
1210 not found

And jcmd with no arguments lists only the JVMs you can see:

$ jcmd
4410 jdk.jcmd/sun.tools.jcmd.JCmd
$ sudo jcmd
1210 /opt/app/orders.jar
4411 jdk.jcmd/sun.tools.jcmd.JCmd

jcmd: one tool for everything

$ sudo -u appuser jcmd $(pgrep -f orders.jar) help
1210:
The following commands are available:
Compiler.CodeHeap_Analytics
...
GC.class_histogram
GC.heap_dump
GC.heap_info
GC.run
JFR.start
Thread.print
VM.command_line
VM.flags
VM.native_memory
VM.system_properties
VM.uptime
VM.version
help

For more information about a specific command use 'help <command>'.

The ones you will use weekly:

jcmd PID VM.version              which JDK build is actually running
jcmd PID VM.command_line         the real flags, including JAVA_TOOL_OPTIONS
jcmd PID VM.flags                every non-default flag, ergonomic ones included
jcmd PID VM.system_properties    -D properties, user.dir, java.io.tmpdir
jcmd PID GC.heap_info            heap committed/used, metaspace
jcmd PID GC.class_histogram      live objects by class (forces a full GC)
jcmd PID GC.heap_dump FILE       HPROF dump for Eclipse MAT
jcmd PID Thread.print [-l]       thread dump
jcmd PID VM.native_memory summary   native memory (needs NMT at start)
jcmd PID JFR.start duration=60s filename=/tmp/rec.jfr   flight recording

You can also target by name instead of PID - handy in scripts:

$ sudo -u appuser jcmd orders.jar VM.uptime
1210:
8226.493 s

VM.flags, decoded

$ sudo -u appuser jcmd $(pgrep -f orders.jar) VM.flags
1210:
-XX:CICompilerCount=2 -XX:ConcGCThreads=1 -XX:G1ConcRefinementThreads=2 ... -XX:G1HeapRegionSize=1048576 ... -XX:InitialHeapSize=96468992 ... -XX:MaxHeapSize=536870912 -XX:MaxNewSize=321912832 ... -XX:ReservedCodeCacheSize=251658240 -XX:+SegmentedCodeCache -XX:SoftMaxHeapSize=536870912 ... -XX:+UseCompressedOops -XX:+UseFastUnorderedTimeStamps -XX:+UseG1GC

What to read out of it: the collector (+UseG1GC), the heap bounds (InitialHeapSize 92 MB, MaxHeapSize 512 MB), NativeMemoryTracking if it is on, HeapDumpOnOutOfMemoryError if it is set. -XX:+PrintFlagsFinal is the same information for a JVM that is not running yet.

The paths are the target's, not yours

GC.heap_dump is performed by the target JVM, as the target's user, with a path relative to the target's working directory:

$ sudo -u appuser jcmd $(pgrep -f orders.jar) GC.heap_dump heap.hprof
1210:
Dumping heap to /opt/app/heap.hprof ...
Heap dump file created [301451712 bytes in 0.793 secs]

$ sudo -u appuser jcmd $(pgrep -f orders.jar) GC.heap_dump /home/learner/heap.hprof
1210:
Dumping heap to /home/learner/heap.hprof ...
Unable to create /home/learner/heap.hprof: Permission denied

Use an absolute path in a directory the JVM's user can write, on a filesystem with room for a file the size of the heap - /tmp on a container is often a small overlay. The dump is created mode 0600, owned by the JVM's user.

jmap -dump:live,format=b,file=heap.hprof PID resolves the path on your side instead. And jmap -heap is gone since JDK 9:

$ sudo jmap -heap $(pgrep -f orders.jar)
Error: -heap option used
Cannot connect to core dump or remote debug server. Use jhsdb jmap instead

kill -3

The oldest trick: kill -3 PID (SIGQUIT) makes the JVM print a thread dump to its own stdout. For a systemd service that is the journal; in Kubernetes it is kubectl logs. No tools needed, but you have to go and find it.

$ sudo kill -3 $(pgrep -f orders.jar)
$ journalctl -u orders --since '-1min' | grep -c 'java.lang.Thread.State'
46

Why it helps

The first time you need a thread dump from a stuck production service is not the moment to find out the image has no jcmd. Knowing this lets you make the decision in advance: a JDK-based runtime image, an ephemeral debug container, or Actuator endpoints. That is a platform-team decision you will be asked about when writing base image standards.

It also saves you time during incidents. Operation not permitted, 1210 not found and Unable to open socket file all look like "the tool is broken", but each means something specific: wrong user, can't read hsperfdata, or a JVM that is not answering. And you will know that a heap dump path is resolved by the target JVM, so the file lands in its working directory, not yours.

Commands in this lesson

jcmd apt jstat jmap kill journalctl

FAQ

What is the difference between jcmd, jstack and jmap?

jcmd is the general tool: it sends any diagnostic command to a running JVM, including Thread.print, GC.heap_dump, GC.class_histogram, VM.flags and VM.native_memory. jstack only takes thread dumps and jmap only does heap histograms and dumps. They use the same attach mechanism. Modern practice is to use jcmd for everything; jstack and jmap are still around and you will see them in older runbooks. jmap -heap no longer exists since JDK 9.

Why does jstat say "not found" when the process clearly exists?

jstat does not attach. It reads the JVM's shared performance counters from /tmp/hsperfdata_<user>/<pid>, a file with mode 0600 owned by the JVM's user. If you are a different user you cannot read it, and jstat reports 1210 not found. It means "I cannot see its counters", not "no such process". Run it as the JVM's user: sudo -u appuser jstat -gcutil 1210.

Is it safe to run these tools against production?

Mostly, with a few exceptions you should know. Flags, heap info, thread dumps and jstat are cheap: a thread dump pauses the JVM for milliseconds. GC.class_histogram and jmap -histo:live force a full GC first, which is a real pause on a large heap. A heap dump pauses the JVM for seconds per GB and writes a file as big as the live heap, full of sensitive data. Choose deliberately.

What does kill -3 do to a Java process?

It sends SIGQUIT, and the JVM does not quit: it prints a thread dump to its own stdout and carries on. For a systemd service that output goes to the journal; in Kubernetes it goes to kubectl logs. It needs no JDK tools, which makes it useful in JRE images, but you have to go and find the dump in the logs, and distroless images often have no kill binary to run it.

Why was my heap dump written to /opt/app instead of my home directory?

jcmd PID GC.heap_dump heap.hprof asks the target JVM to write the dump, so the relative path is resolved in the JVM's working directory and the file is written as the JVM's user. That is why it landed in /opt/app, and why an absolute path in your home directory fails with Permission denied: appuser cannot write there. Use an absolute path the JVM's user can write, on a filesystem with room for the whole heap.

In an interview Mid

How do you take a thread dump of a Java process, and why might it fail?

jcmd, jstack and jmap use the attach mechanism (a SIGQUIT plus a UNIX socket in the JVM's /tmp), which is why they fail:

Also asked: What is the difference between the JRE and the JDK, and why does it matter on servers? · Where does a heap dump taken with jcmd get written? · Why does jstat say "not found" for a process that is running?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.