Clusters, not individuals
The problem. A dump of 200 threads is unreadable one by one. Five patterns cover almost every stuck Java service; once you can recognise them, the dump names the culprit in minutes.
What you need to know already: thread-dump format, states and lock lines (20.20), connection pools (17.22), timeouts (9.1).
A thread pool (or executor) is a fixed set of worker threads that take jobs from a queue - the web server's request threads are one. A deadlock is two threads each holding a lock the other one needs, waiting forever.
Do not read a dump top to bottom. Count. Fifty threads with the same stack is the answer; one odd stack almost never is. The triage grep from the last lesson, on a sick orders service:
# dump1.txt = a dump of a sick orders service (not on this box)
grep -A1 'java.lang.Thread.State' dump1.txt | grep -v -- '--' | sort | uniq -c | sort -rn | head
192 java.lang.Thread.State: TIMED_WAITING (parking)
190 at jdk.internal.misc.Unsafe.park([email protected]/Native Method)
21 java.lang.Thread.State: RUNNABLE
10 at sun.nio.ch.Net.poll([email protected]/Native Method)
2 java.lang.Thread.State: WAITING (parking)
park is too generic to be the answer. Look a few frames down in one of them:
"http-nio-8080-exec-131" #188 ... waiting on condition
java.lang.Thread.State: TIMED_WAITING (parking)
at jdk.internal.misc.Unsafe.park([email protected]/Native Method)
- parking to wait for <0x00000000e94b9b80> (a java.util.concurrent.SynchronousQueue$Transferer)
at java.util.concurrent.locks.LockSupport.parkNanos([email protected]/LockSupport.java:269)
...
at com.zaxxer.hikari.util.ConcurrentBag.borrow(ConcurrentBag.java:151)
at com.zaxxer.hikari.pool.HikariPool.getConnection(HikariPool.java:162)
...
at lab.orders.web.OrdersController.list(OrdersController.java:29)
And count that frame:
# the same sick dump1.txt
grep -c 'HikariPool.getConnection(HikariPool.java:162)' dump1.txt
180
grep -c 'PaymentsClient.authorize' dump1.txt
20
Pattern 1: connection pool exhaustion
180 request threads are waiting to borrow a database connection (TIMED_WAITING because Hikari waits up to connectionTimeout, 30 s by default). The pool has 20 connections. Where are they? The 20 threads in PaymentsClient.authorize - each is inside a DB transaction, holding a connection, while it waits for the payments service to answer. And:
http.client.read.timeout.ms=0
No read timeout. The big cluster is the symptom; the small cluster holding the resource is the cause. From outside it presents as latency, not errors: requests wait up to 30 s for a connection and then succeed - p99 goes vertical, the error rate stays flat, CPU is idle.
The cascade from there: every endpoint that needs the database is now slow, not just checkout. Tomcat's 200 threads fill up. Health checks that touch the database start timing out. An orchestrator restarts the pod; its traffic moves to the remaining pods; they fill up the same way. One slow dependency took the service down, because nothing bounded how long a thread could wait.
Pattern 2: the downstream hang
All request threads RUNNABLE in a socket read to the same client class:
java.lang.Thread.State: RUNNABLE
at sun.nio.ch.Net.poll([email protected]/Native Method)
at sun.nio.ch.NioSocketImpl.park(...)
...
at org.apache.hc.core5.http.impl.io.SessionInputBufferImpl.fillBuffer(SessionInputBufferImpl.java:149)
...
at lab.orders.payments.PaymentsClient.authorize(PaymentsClient.java:48)
The process is alive, CPU near zero, nothing in the logs, and nothing will throw: without a read timeout a blocking read waits forever. The same threads, with the same stacks, in three dumps taken ten seconds apart, with cpu= not moving. The fix is a read timeout (and then a circuit breaker - next chapter).
Pattern 3: deadlock
The JVM finds monitor deadlocks for you, at the end of the dump:
Found one Java-level deadlock:
=============================
"settlement-1":
waiting to lock monitor 0x0000ffff5c004e80 (object 0x00000000f5a1b2c8, a lab.payments.settlement.Ledger),
which is held by "settlement-2"
"settlement-2":
waiting to lock monitor 0x0000ffff5c006f10 (object 0x00000000f5a1c4e0, a lab.payments.settlement.AccountBook),
which is held by "settlement-1"
Java stack information for the threads listed above:
===================================================
"settlement-1":
at lab.payments.settlement.Ledger.post(Ledger.java:42)
- waiting to lock <0x00000000f5a1b2c8> (a lab.payments.settlement.Ledger)
at lab.payments.settlement.AccountBook.transfer(AccountBook.java:88)
- locked <0x00000000f5a1c4e0> (a lab.payments.settlement.AccountBook)
...
Found 1 deadlock.
Two threads each hold what the other wants: settlement-1 holds the AccountBook and wants the Ledger; settlement-2 holds the Ledger and wants the AccountBook. They will wait forever. Both show BLOCKED (on object monitor).
What you do: capture the dump (it is the bug report), restart to restore service, and hand the two stacks to the developers - the fix is a consistent lock order. Rare in practice, overrepresented in interviews, and the health endpoint stays UP through the whole thing, because nothing it checks is deadlocked.
Deadlocks on java.util.concurrent locks (ReentrantLock) are also detected, with -l. A deadlock between a thread and a database row lock is not - that shows up as threads RUNNABLE in a socket read, waiting for the database to answer.
Pattern 4: async pool exhaustion
With CompletableFuture the blocking moves off the request threads into an executor. When that executor is exhausted nothing throws: tasks queue in an unbounded LinkedBlockingQueue and requests silently stop progressing.
"orders-async-1" ... RUNNABLE Net.poll ... ShippingClient.eta ... CompletableFuture$AsyncSupply.run
"orders-async-2" ... RUNNABLE (same)
...
"orders-async-8" ... RUNNABLE (same)
"http-nio-8080-exec-1..10" ... WAITING (parking) ... TaskQueue.take <- Tomcat is idle!
All 8 executor threads stuck in the same downstream call; the Tomcat threads look idle, because async request handling released them. The queue depth is in the executor_queued_tasks metric, not in the dump. Two things make this worse on small boxes:
CompletableFuture.supplyAsync(fn)without an executor usesForkJoinPool.commonPool(), sized CPUs - 1. On a 2-CPU pod that is 1, and when parallelism is 1 or less the JDK does not use the pool at all - it starts a new thread per task (ThreadPerTaskExecutor). Either way you did not choose the concurrency. Always pass your own, named, bounded executor.- A bounded pool with an unbounded queue is only half bounded. Bound the queue too, and decide what happens when it is full (reject fast - backpressure).
Pattern 5: thread leak
# dump.txt = a dump of a service leaking threads (not on this box)
grep '^"' dump.txt | sed -E 's/"([^"]*[^0-9-])[0-9-]*".*/\1/' | sort | uniq -c | sort -rn | head -3
412 pool-
10 http-nio-8080-exec-
2 settlement-
Hundreds of pool-N-thread-1 threads - N climbing - every one idle in ThreadPoolExecutor.getTask. Code that creates Executors.newFixedThreadPool per request and never shuts it down. Each thread costs stack and native memory, so it ends as a cgroup OOM kill. jcmd PID VM.native_memory summary.diff shows the Thread category climbing (this chapter's incident).
Three dumps, ten seconds apart
One dump is a photograph: you cannot tell stuck from busy. Three tell you:
for i in 1 2 3; do sudo -u appuser jcmd PID Thread.print > dump$i.txt; sleep 10; done
same thread, same stack, same cpu= in all three stuck
same stack cluster, different threads each time busy but moving (a hot path)
cluster grows from 1 to 3 getting worse right now
In Kubernetes: for i in 1 2 3; do kubectl exec pod -- jcmd 1 Thread.print > dump$i.txt; sleep 10; done.
When a dump is not enough: JFR and async-profiler
- JFR (Java Flight Recorder), built into the JVM, low overhead, safe in production:
jcmd PID JFR.start duration=60s filename=/tmp/app.jfr. Records CPU samples, allocation, locks, GC, I/O, and which threads waited on what and for how long. Open the file in JDK Mission Control. - async-profiler - an external agent that produces flame graphs of CPU, allocations or locks. The fastest route to "which method is burning the CPU".
A dump answers "what is everyone doing now"; a recording answers "where did the last minute go".