OnCallReady

Lesson 20.22 · JVM Internals · 18 min read

Thread dump patterns: clusters, exhaustion, deadlock, async

In plain words

Imagine a supermarket where the queue at the tills stretches out of the door. You could interview each customer, or you could step back and count: 180 people are waiting, and there are 20 tills, and every cashier is on the phone to the manager, who is not answering. The long queue is the symptom; the 20 cashiers stuck on the phone are the cause.

Reading thread dumps is the same: count the clusters instead of reading every thread. In orders, 180 threads wait in HikariPool.getConnection and 20 sit in PaymentsClient.authorize holding the connections, because http.client.read.timeout.ms=0 means they wait forever. Deadlocks, async pool exhaustion and thread leaks each have their own shape in the count.

Clusters, not individuals

The problem. A dump of 200 threads is unreadable one by one. Five patterns cover almost every stuck Java service; once you can recognise them, the dump names the culprit in minutes.

What you need to know already: thread-dump format, states and lock lines (20.20), connection pools (17.22), timeouts (9.1).

A thread pool (or executor) is a fixed set of worker threads that take jobs from a queue - the web server's request threads are one. A deadlock is two threads each holding a lock the other one needs, waiting forever.

Do not read a dump top to bottom. Count. Fifty threads with the same stack is the answer; one odd stack almost never is. The triage grep from the last lesson, on a sick orders service:

# dump1.txt = a dump of a sick orders service (not on this box)
grep -A1 'java.lang.Thread.State' dump1.txt | grep -v -- '--' | sort | uniq -c | sort -rn | head
    192    java.lang.Thread.State: TIMED_WAITING (parking)
    190 	at jdk.internal.misc.Unsafe.park([email protected]/Native Method)
     21    java.lang.Thread.State: RUNNABLE
     10 	at sun.nio.ch.Net.poll([email protected]/Native Method)
      2    java.lang.Thread.State: WAITING (parking)

park is too generic to be the answer. Look a few frames down in one of them:

"http-nio-8080-exec-131" #188 ... waiting on condition
   java.lang.Thread.State: TIMED_WAITING (parking)
	at jdk.internal.misc.Unsafe.park([email protected]/Native Method)
	- parking to wait for  <0x00000000e94b9b80> (a java.util.concurrent.SynchronousQueue$Transferer)
	at java.util.concurrent.locks.LockSupport.parkNanos([email protected]/LockSupport.java:269)
	...
	at com.zaxxer.hikari.util.ConcurrentBag.borrow(ConcurrentBag.java:151)
	at com.zaxxer.hikari.pool.HikariPool.getConnection(HikariPool.java:162)
	...
	at lab.orders.web.OrdersController.list(OrdersController.java:29)

And count that frame:

# the same sick dump1.txt
grep -c 'HikariPool.getConnection(HikariPool.java:162)' dump1.txt
180
grep -c 'PaymentsClient.authorize' dump1.txt
20

Pattern 1: connection pool exhaustion

180 request threads are waiting to borrow a database connection (TIMED_WAITING because Hikari waits up to connectionTimeout, 30 s by default). The pool has 20 connections. Where are they? The 20 threads in PaymentsClient.authorize - each is inside a DB transaction, holding a connection, while it waits for the payments service to answer. And:

http.client.read.timeout.ms=0

No read timeout. The big cluster is the symptom; the small cluster holding the resource is the cause. From outside it presents as latency, not errors: requests wait up to 30 s for a connection and then succeed - p99 goes vertical, the error rate stays flat, CPU is idle.

The cascade from there: every endpoint that needs the database is now slow, not just checkout. Tomcat's 200 threads fill up. Health checks that touch the database start timing out. An orchestrator restarts the pod; its traffic moves to the remaining pods; they fill up the same way. One slow dependency took the service down, because nothing bounded how long a thread could wait.

Pattern 2: the downstream hang

All request threads RUNNABLE in a socket read to the same client class:

   java.lang.Thread.State: RUNNABLE
	at sun.nio.ch.Net.poll([email protected]/Native Method)
	at sun.nio.ch.NioSocketImpl.park(...)
	...
	at org.apache.hc.core5.http.impl.io.SessionInputBufferImpl.fillBuffer(SessionInputBufferImpl.java:149)
	...
	at lab.orders.payments.PaymentsClient.authorize(PaymentsClient.java:48)

The process is alive, CPU near zero, nothing in the logs, and nothing will throw: without a read timeout a blocking read waits forever. The same threads, with the same stacks, in three dumps taken ten seconds apart, with cpu= not moving. The fix is a read timeout (and then a circuit breaker - next chapter).

Pattern 3: deadlock

The JVM finds monitor deadlocks for you, at the end of the dump:

Found one Java-level deadlock:
=============================
"settlement-1":
  waiting to lock monitor 0x0000ffff5c004e80 (object 0x00000000f5a1b2c8, a lab.payments.settlement.Ledger),
  which is held by "settlement-2"

"settlement-2":
  waiting to lock monitor 0x0000ffff5c006f10 (object 0x00000000f5a1c4e0, a lab.payments.settlement.AccountBook),
  which is held by "settlement-1"

Java stack information for the threads listed above:
===================================================
"settlement-1":
	at lab.payments.settlement.Ledger.post(Ledger.java:42)
	- waiting to lock <0x00000000f5a1b2c8> (a lab.payments.settlement.Ledger)
	at lab.payments.settlement.AccountBook.transfer(AccountBook.java:88)
	- locked <0x00000000f5a1c4e0> (a lab.payments.settlement.AccountBook)
	...
Found 1 deadlock.

Two threads each hold what the other wants: settlement-1 holds the AccountBook and wants the Ledger; settlement-2 holds the Ledger and wants the AccountBook. They will wait forever. Both show BLOCKED (on object monitor).

What you do: capture the dump (it is the bug report), restart to restore service, and hand the two stacks to the developers - the fix is a consistent lock order. Rare in practice, overrepresented in interviews, and the health endpoint stays UP through the whole thing, because nothing it checks is deadlocked.

Deadlocks on java.util.concurrent locks (ReentrantLock) are also detected, with -l. A deadlock between a thread and a database row lock is not - that shows up as threads RUNNABLE in a socket read, waiting for the database to answer.

Pattern 4: async pool exhaustion

With CompletableFuture the blocking moves off the request threads into an executor. When that executor is exhausted nothing throws: tasks queue in an unbounded LinkedBlockingQueue and requests silently stop progressing.

"orders-async-1" ... RUNNABLE   Net.poll ... ShippingClient.eta ... CompletableFuture$AsyncSupply.run
"orders-async-2" ... RUNNABLE   (same)
...
"orders-async-8" ... RUNNABLE   (same)
"http-nio-8080-exec-1..10" ... WAITING (parking) ... TaskQueue.take     <- Tomcat is idle!

All 8 executor threads stuck in the same downstream call; the Tomcat threads look idle, because async request handling released them. The queue depth is in the executor_queued_tasks metric, not in the dump. Two things make this worse on small boxes:

Pattern 5: thread leak

# dump.txt = a dump of a service leaking threads (not on this box)
grep '^"' dump.txt | sed -E 's/"([^"]*[^0-9-])[0-9-]*".*/\1/' | sort | uniq -c | sort -rn | head -3
    412 pool-
     10 http-nio-8080-exec-
      2 settlement-

Hundreds of pool-N-thread-1 threads - N climbing - every one idle in ThreadPoolExecutor.getTask. Code that creates Executors.newFixedThreadPool per request and never shuts it down. Each thread costs stack and native memory, so it ends as a cgroup OOM kill. jcmd PID VM.native_memory summary.diff shows the Thread category climbing (this chapter's incident).

Three dumps, ten seconds apart

One dump is a photograph: you cannot tell stuck from busy. Three tell you:

for i in 1 2 3; do sudo -u appuser jcmd PID Thread.print > dump$i.txt; sleep 10; done
same thread, same stack, same cpu= in all three      stuck
same stack cluster, different threads each time     busy but moving (a hot path)
cluster grows from 1 to 3                           getting worse right now

In Kubernetes: for i in 1 2 3; do kubectl exec pod -- jcmd 1 Thread.print > dump$i.txt; sleep 10; done.

When a dump is not enough: JFR and async-profiler

A dump answers "what is everyone doing now"; a recording answers "where did the last minute go".

Why it helps

These five patterns cover most "the service is up but nothing works" incidents you will see with Java: pool exhaustion, a hung downstream, a deadlock, a stuck async executor and a thread leak. Recognising the shape turns a two-hour war room into a ten-minute diagnosis with a clear owner.

The biggest lesson is that the big cluster is rarely the cause. In an incident, everyone looks at the 180 threads waiting for the database and blames the database. You look for the small cluster holding the resource, find the missing read timeout, and the fix is a config change instead of a bigger database. It is also the foundation for the next chapter's timeouts and circuit breakers, which exist to prevent exactly these patterns.

FAQ

Why does pool exhaustion show up as latency instead of errors?

Because waiting for a connection is not an error until the wait times out. Hikari waits up to connectionTimeout, 30 seconds by default, and many requests get a connection just before that and succeed. So p99 shoots up, error rate stays flat and CPU is idle. Only when waits exceed the timeout do you see Connection is not available, request timed out. By then the Tomcat pool is full and everything is slow, not just the endpoint that caused it.

Why is the health endpoint UP during a deadlock?

Because nothing the health check touches is deadlocked. Two settlement threads waiting on each other's locks do not affect the Tomcat thread that answers /actuator/health, and the database check still gets a connection. So the service looks healthy while one function is permanently stuck. This is why liveness probes do not catch every hang, and why a thread dump and business metrics like "settlements completed per minute" matter.

What should I do when I find a deadlock?

Capture the evidence first: the thread dump, which the JVM annotates with Found one Java-level deadlock and both stacks. Then restart the service to restore it, because deadlocked threads never recover by themselves. Hand the two stacks to the developers: the fix is almost always a consistent lock order in the code. Note that a deadlock with a database row lock is not detected by the JVM; it looks like threads RUNNABLE in a socket read.

What is special about CompletableFuture's default executor?

supplyAsync without an executor uses ForkJoinPool.commonPool(), sized at CPUs minus 1. On a 2-CPU pod that is 1, and when parallelism is 1 or less the JDK does not use the pool at all: it starts a new thread per task. Either way, you did not choose the concurrency, and blocking calls in it starve other users of the common pool. Always pass your own named, bounded executor, with a bounded queue.

How do I spot a thread leak in a dump?

Count threads per pool name: grep '^"' dump.txt | sed -E 's/"([^"]*[^0-9-])[0-9-]*".*/\1/' | sort | uniq -c | sort -rn. A leak shows as hundreds of pool-N-thread-1 with N climbing, each idle in ThreadPoolExecutor.getTask. That is code creating a new executor per request and never shutting it down. It ends as a cgroup OOM kill, because every thread costs stack and native memory, while the heap graph stays flat.

In an interview Mid

Requests to a Java service hang, CPU is low and there are no errors. How do you find out why?

Take three thread dumps ten seconds apart (jcmd PID Thread.print > dump$i.txt) and count clusters of identical stacks. Same thread, same stack, same cpu= in all three = stuck. Then match the pattern:

Fix is usually a timeout; for a deadlock, keep the dump, restart, hand the stacks to the developers.

Also asked: What is a deadlock and how do you detect one in Java? · Why take three thread dumps instead of one? · How can one slow dependency take down a whole service?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.