OnCallReady

Lesson 20.20 · JVM Internals · 14 min read

Reading a thread dump: format and states

In plain words

Imagine taking a photo of a busy kitchen at one exact moment. Everyone is caught mid-action: one cook is stirring, three are standing at the fridge waiting for the one cook who is inside it, another is on the phone waiting for a supplier to answer. Each has a name badge. From one photo you can see who is doing what and who is waiting for whom.

A thread dump is that photo of a JVM. Every thread has a name like http-nio-8080-exec-7, a state like RUNNABLE or WAITING, and a stack showing what it is doing, most recent first. jcmd PID Thread.print takes it. The trap: a thread waiting for a network reply shows as RUNNABLE, like the cook on the phone who looks busy but is just waiting.

Taking one

The problem. A service stops answering but uses no CPU. Nothing is in the logs. A thread dump - a snapshot of what every thread in the JVM is doing right now - is how you see where it is stuck.

What you need to know already: threads (20.1), kill -3 / SIGQUIT (3.6), jcmd PID Thread.print (20.3), grep and sort | uniq -c (7.1, 7.4).

A stack trace is the list of method calls a thread is inside, innermost first (a method = a function belonging to a class). A lock (or monitor) is a guard only one thread can hold at a time; others that want it wait.

sudo -u appuser jcmd PID Thread.print > dump1.txt       preferred
sudo -u appuser jcmd PID Thread.print -l > dump1.txt    with java.util.concurrent locks
sudo -u appuser jstack -l PID > dump1.txt               the older tool, same content, no "PID:" header
sudo kill -3 PID                                        to the JVM's stdout (the journal)
curl -H 'Accept: text/plain' localhost:8080/actuator/threaddump    over HTTP (next chapter)

Taking a dump pauses the JVM at a safepoint for a few milliseconds. Safe in production.

The header

1210:
2026-09-23 10:15:08
Full thread dump OpenJDK 64-Bit Server VM (21.0.8+9-Ubuntu-0ubuntu1~26.04 mixed mode, sharing):

Threads class SMR info:
_java_thread_list=0x0000ffff5c002e10, length=46, elements={
0x0000ffff943d7320, 0x0000ffff94043210, 0x0000ffff943f6af0, 0x0000ffff94023930,
...
}

length=46 is the number of Java threads. The SMR block is internal bookkeeping - skip it.

One thread

"http-nio-8080-exec-7" #52 [1297] daemon prio=5 os_prio=0 cpu=95.10ms elapsed=3590.11s tid=0x0000ffff9410e970 nid=1297 runnable  [0x0000ffff7b4b8000]
   java.lang.Thread.State: RUNNABLE
	at sun.nio.ch.Net.poll([email protected]/Native Method)
	at sun.nio.ch.NioSocketImpl.park([email protected]/NioSocketImpl.java:191)
	at sun.nio.ch.NioSocketImpl.implRead([email protected]/NioSocketImpl.java:309)
	...
	at org.springframework.web.client.RestTemplate.postForObject(RestTemplate.java:519)
	at lab.orders.payments.PaymentsClient.authorize(PaymentsClient.java:48)
	at lab.orders.checkout.CheckoutService.checkout(CheckoutService.java:71)
	...
	at java.lang.Thread.run([email protected]/Thread.java:1583)

   Locked ownable synchronizers:
	- <0x00000000f1a2b3c8> (a org.apache.tomcat.util.threads.ThreadPoolExecutor$Worker)

Field by field:

The four states that matter

RUNNABLE                   executing, or blocked in native code - including a socket read
BLOCKED (on object monitor)   waiting to enter a synchronized block another thread holds
WAITING (parking)          LockSupport.park with no timeout: a pool, a queue, a future.get()
WAITING (on object monitor)   Object.wait() with no timeout
TIMED_WAITING (parking)    park with a timeout: pool borrow with timeout, poll(timeout)
TIMED_WAITING (sleeping)   Thread.sleep()

The trap in the list: a thread waiting for bytes on a socket is RUNNABLE. From the JVM's point of view it is inside a native call (Net.poll, socketRead0 on older JDKs) and could return any moment. So "all threads RUNNABLE" does not mean "all threads busy on CPU": check the top frames. The cpu= field across two dumps settles it.

RUNNABLE + top frame Net.poll / socketRead / EPoll.wait    waiting on the network
RUNNABLE + top frame in your code, cpu= growing            actually computing

The lock lines

- parking to wait for  <0x00000000e94b9b80> (a java.util.concurrent.SynchronousQueue$Transferer)
- waiting to lock <0x00000000f5a1b2c8> (a lab.payments.settlement.Ledger)
- locked <0x00000000f5a1c4e0> (a lab.payments.settlement.AccountBook)
- waiting on <0x...> (a java.lang.Object)

The hex number is the object's identity. Same address in many threads = they all wait for the same thing. Search the dump for who locked it.

Threads you can ignore

Every JVM has them, and they look scary if you do not know them:

"Reference Handler", "Finalizer", "Signal Dispatcher", "Service Thread",
"Monitor Deflation Thread", "C1/C2 CompilerThread0", "Common-Cleaner",
"Notification Thread", "Attach Listener" (appears once you attach!),
"DestroyJavaVM", and the non-Java ones at the end:
"VM Thread", "GC Thread#0", "G1 Main Marker", "G1 Conc#0", "G1 Refine#0", "VM Periodic Task Thread"

And the dump ends with JNI global refs: 24, weak refs: 0 - and, when there is one, the deadlock report.

Triage greps

grep -c 'java.lang.Thread.State' dump1.txt                  how many Java threads
grep 'java.lang.Thread.State' dump1.txt | sort | uniq -c | sort -rn
                                                            threads per state
grep -A1 'java.lang.Thread.State' dump1.txt | grep -v -- '--' | sort | uniq -c | sort -rn | head
                                                            state + top frame: the clusters
grep '^"' dump1.txt | sed -E 's/"([^"]*[^0-9-])[0-9-]*".*/\1/' | sort | uniq -c | sort -rn
                                                            threads per pool name
grep -B2 -A20 '"http-nio-8080-exec-7"' dump1.txt            one thread in full

The next lesson turns those counts into diagnoses.

Why it helps

A thread dump is the single most useful artifact when a Java service is alive but not answering: no errors in the logs, CPU idle, health check green, and requests hanging. Knowing how to read it lets you say in minutes "180 request threads are waiting for a DB connection" or "every thread is stuck reading from payments", which tells the team where to look.

It is also where you catch surprises: an anonymous pool-N-thread-M executor someone forgot to name, a thread count far higher than expected, or Attach Listener showing someone already attached. And when developers ask for "logs" from a hung service, a dump plus the triage greps is a far better answer.

FAQ

Why is a thread waiting on a socket shown as RUNNABLE?

Because from the JVM's point of view it is inside a native call, Net.poll or socketRead0, and could return at any moment. The JVM does not know that the kernel has put the thread to sleep waiting for bytes. So "all threads RUNNABLE" does not mean "all threads on CPU". Check the top frame: a network call means waiting. Compare the cpu= field across two dumps: if it does not move, the thread is not computing.

What is the difference between BLOCKED and WAITING?

BLOCKED means the thread wants to enter a synchronized block that another thread holds; it shows waiting to lock <address>. WAITING means it parked itself deliberately and waits to be woken: a pool, a queue, Future.get(), Object.wait(), all without a timeout. TIMED_WAITING is the same with a timeout, such as a Hikari connection borrow with connectionTimeout or Thread.sleep. Many BLOCKED threads on one address means lock contention.

How do I match a thread in top -H with a thread dump?

top -H -p PID shows the OS thread id in the PID column. In a dump from a modern JDK, that same number appears in brackets after the Java id, #52 [1297], and as nid=1297. Older JDKs printed nid in hex, so you needed printf '%x\n' 1297. The cpu= field in the dump then tells you how much that thread has consumed in total.

Which threads in a dump can I ignore?

The JVM's own housekeeping threads: Reference Handler, Finalizer, Signal Dispatcher, Service Thread, Monitor Deflation Thread, C1/C2 CompilerThread, Common-Cleaner, Notification Thread, Attach Listener (created because you attached), DestroyJavaVM, and the non-Java ones at the end such as VM Thread, GC Thread#0 and G1 Conc#0. They exist in every JVM. Focus on the application's pools, identified by name.

Does taking a thread dump hurt production?

Very little. The JVM stops at a safepoint for a few milliseconds while it walks the stacks, then continues. It is safe to take several in a row, and three dumps about ten seconds apart is the standard practice. The only risk is on a JVM that is already stuck in a long GC: the dump waits for the safepoint too, and the attach can time out.

In an interview Mid

What does a Java thread dump contain, and what do you look at first?

A snapshot of every thread: for each one the name, Java id, nid (the OS thread id top -H shows), cpu= used so far, the java.lang.Thread.State, the stack (most recent frame first) and lock lines (waiting to lock <0x...>, locked <0x...>). It pauses the JVM for milliseconds - safe in production.

What I look at first:

Also asked: Explain the Java thread states and what each one suggests during an incident. · Is taking a thread dump safe on a production service? · How do you match a thread that is hot in top -H to a thread in a dump?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.