In plain words
Imagine a restaurant where waiters take orders and hand them to cooks in the back instead of cooking themselves. The waiters are always free, so the front of house looks calm. But if all eight cooks are stuck waiting for a delivery, orders pile up on the kitchen spike, and customers just wait. The waiters look idle, the queue grows out of sight, and nobody shouts that anything is wrong.
That is async request handling with CompletableFuture. Tomcat's request threads are the waiters and are released immediately; the work runs on an executor, the cooks. When the executor's threads are all blocked on a slow call, tasks queue silently in an unbounded queue. Nothing throws. The fix is a named, bounded executor with a bounded queue and a timeout inside each task.
Moving the blocking somewhere else
The problem. Teams make a slow endpoint "async" to free Tomcat threads. Then one day the endpoint spins forever with no errors, Tomcat is idle, and nothing in the logs. The waiting did not go away; it moved to a pool nobody watches.
What you need to know already: thread dumps and their patterns (20.20, 20.22), pools and bulkheads (21.15, 21.22), the metrics page (21.11). From JavaScript: a Promise.
A controller is the Java method that handles one URL. An async controller returns a CompletableFuture (Java's version of a JS Promise: a value that will arrive later; or DeferredResult, or a reactive type). Tomcat's request thread is released immediately; the work runs on an executor (a thread pool that runs submitted tasks from a queue); when the future completes, Spring writes the response: (or DeferredResult, or a reactive type). Tomcat's request thread is released immediately; the work runs on an executor; when the future completes, Spring writes the response:
@GetMapping("/api/orders/{id}/eta")
public CompletableFuture<Eta> eta(@PathVariable long id) {
return CompletableFuture.supplyAsync(() -> shipping.eta(id), ordersAsync);
}
This frees Tomcat threads - and hides the problem when shipping.eta blocks: the executor's threads fill up instead, and extra tasks wait in its queue.
What exhaustion looks like
thread dump
"orders-async-1..8" RUNNABLE Net.poll ... ShippingClient.eta ... CompletableFuture$AsyncSupply.run
"http-nio-8080-exec-1..10" WAITING TaskQueue.take <- Tomcat looks idle
metrics
executor_active_threads{name="ordersAsync"} 8.0 = pool size
executor_queued_tasks{name="ordersAsync"} 254.0 and growing
http_server_requests_seconds_max{uri="/api/orders/{id}/eta"} 30.0
Nothing throws. Requests wait in the queue until Spring MVC's async request timeout (spring.mvc.async.request-timeout; on Tomcat 30 s by default) answers 503 with AsyncRequestTimeoutException - and the task is still in the queue, so it will run later for a client that has gone. With ThreadPoolExecutor's default unbounded LinkedBlockingQueue, the queue grows without limit.
The common pool trap
CompletableFuture.supplyAsync(() -> slowCall()) // no executor argument
uses ForkJoinPool.commonPool() (the JVM's one built-in shared pool): shared by every caller in the JVM (parallel streams too), sized CPUs - 1 with a minimum of 1. When that parallelism is 1 - any machine or pod with 2 CPUs or fewer - CompletableFuture does not use the common pool at all: it starts a new thread for every task. So on a 2-CPU pod, supplyAsync without an executor gives you unbounded threads; on a 16-CPU node, 15 shared threads that one blocking call can monopolise. Neither was a decision.
Rule: always pass a named, bounded executor for anything that blocks, one per dependency (a bulkhead):
@Bean ThreadPoolTaskExecutor ordersAsync() {
var e = new ThreadPoolTaskExecutor();
e.setThreadNamePrefix("orders-async-"); // readable in dumps
e.setCorePoolSize(8); e.setMaxPoolSize(8);
e.setQueueCapacity(100); // BOUNDED: excess is rejected, not queued forever
return e; // default policy: AbortPolicy -> TaskRejectedException -> 503
}
- Name the threads so a dump says whose pool is stuck.
- Bound the queue so saturation turns into fast rejections (backpressure) instead of 30-second waits.
- Time out the blocking call inside the task (a read timeout, or
orTimeout(2, SECONDS) on the future) so threads come back. - Export the executor's metrics (
executor_active_threads, executor_queued_tasks) and alert on the queue.
Why this matters in a bank-scale platform
A fully asynchronous service has many executors, and the failure always looks the same from outside: latency, no errors, CPU idle, Tomcat idle. The thread dump tells you which executor; its metrics tell you how deep the queue is; the fix is a timeout and a bound.
What you can now do
- Spot async executor exhaustion: executor threads stuck, Tomcat idle, queue growing, no errors.
- Explain the common-pool trap and always pass a named, bounded executor.
- Pick the three fixes: a timeout inside the task, a bounded queue, metrics on the executor.
Why it helps
Fully asynchronous services fail in a way that confuses everyone during incidents: latency is terrible, there are no errors, CPU is idle and even Tomcat looks idle. If you know the pattern, you go straight to the executor's threads in a thread dump and to executor_queued_tasks in the metrics, and you find the downstream call that's stuck.
It also explains a classic small-pod surprise: supplyAsync without an executor on a 2-CPU pod creates a new thread per task, which then shows up as a thread leak or an OOM kill. When you review code or Deployments for async services, you'll know to ask for named executors, bounded queues, timeouts and executor metrics on the dashboard.
FAQ
Why doesn't anything throw when the executor is exhausted?
Because ThreadPoolExecutor's default queue, a LinkedBlockingQueue without a capacity, accepts every task. Nothing is rejected, so nothing fails; tasks wait their turn. Requests only fail when Spring MVC's async request timeout expires, 30 s by default on Tomcat, returning a 503 with AsyncRequestTimeoutException. Even then the task stays queued and runs later for a client that has gone away. A bounded queue turns that into immediate rejections.
What is wrong with supplyAsync without an executor?
It uses ForkJoinPool.commonPool(), which is shared by every caller in the JVM, including parallel streams, and sized at CPUs minus 1. On a 16-CPU node one blocking call can monopolise those 15 shared threads. On a pod with 2 CPUs or fewer, parallelism is 1 and CompletableFuture doesn't use the pool at all: it starts a new thread per task, which is unbounded. Neither is a decision you made.
Why name executor threads?
Because thread dumps and metrics are read by name. orders-async-1 to orders-async-8 all stuck in ShippingClient.eta tells you immediately which pool and which dependency. pool-3-thread-1 tells you nothing, and hundreds of pool-N-thread-1 with N climbing is the signature of a thread leak. Named executors also get their own tags in executor_active_threads and executor_queued_tasks, so you can alert per pool.
Is async faster than blocking?
Not per request. The work takes the same time; async only frees the request thread while it waits. That helps when you have many slow calls and limited threads, because Tomcat's threads can serve other requests. But the executor still has a finite number of threads, and if the downstream call blocks, those threads fill up just as Tomcat's would have. Async moves the bottleneck; it doesn't remove it. Timeouts and bounds still decide what happens under failure.
What does orTimeout do on a CompletableFuture?
future.orTimeout(2, SECONDS) completes the future exceptionally with a TimeoutException if it hasn't finished in time, so the response returns quickly. It does not stop the underlying work: the executor thread keeps running the blocking call. So you still need a read timeout on the client inside the task to actually free the thread. Use both: the future's timeout for the caller, and the client's timeout for the executor.
In an interview Mid
An async service has high latency, no errors and idle Tomcat threads. How do you investigate?
With a CompletableFuture controller the waiting did not go away - it moved from Tomcat's threads to an executor. So I look at the executors:
- Thread dump: the executor's threads (
orders-async-1..8) all RUNNABLE in Net.poll inside the same client call (ShippingClient.eta), while http-nio-8080-exec-* wait in TaskQueue.take - Tomcat looks idle. - Metrics:
executor_active_threads at the maximum and executor_queued_tasks growing. Nothing throws: tasks queue in an unbounded LinkedBlockingQueue until the async request timeout answers 503 - and the task still runs later for a client that has gone.
Then check the code for supplyAsync without an executor: it uses ForkJoinPool.commonPool() (CPUs - 1, shared by the whole JVM), and on 2 CPUs or fewer a new thread per task. Nobody chose the concurrency.
Fixes: a named, bounded executor per dependency, a bounded queue (reject fast), a timeout inside the task (orTimeout or a read timeout), and alerts on the queue metric.
Also asked: What is the risk of CompletableFuture.supplyAsync without an executor argument? · Why is an unbounded queue in front of a thread pool dangerous? · What is the difference between synchronous and asynchronous request handling in a web service?