What a pool does
The problem. A service slows to 25-second responses with zero errors and idle CPU. More often than not a pool is full: all its connections are lent out and every new request waits in line. It looks like "the database is slow" and it is not.
What you need to know already: TCP handshakes and TLS (9.1, 9.15), threads and thread dumps (20.20, 20.22), the metrics page (21.11), latency percentiles (0.2).
Opening a database connection costs a TCP handshake, TLS, authentication and a server process or thread: tens of milliseconds. A pool keeps N connections open and lends them to threads. A thread that wants one when all N are lent out waits.
HikariCP settings that matter
spring.datasource.hikari.maximum-pool-size=10 N. Default 10.
spring.datasource.hikari.minimum-idle=10 default = maximum-pool-size (a fixed pool - recommended)
spring.datasource.hikari.connection-timeout=30000 ms a thread waits to BORROW before it gets an exception. Default 30 s.
spring.datasource.hikari.max-lifetime=1800000 retire connections before the DB or a firewall kills them. Default 30 min.
spring.datasource.hikari.idle-timeout=600000 only matters when minimum-idle < max
spring.datasource.hikari.leak-detection-threshold=20000 log a stack trace if a connection is held > 20 s
spring.datasource.hikari.keepalive-time=120000 ping idle connections (NAT / LB idle timeouts)
When a borrow times out:
java.sql.SQLTransientConnectionException: HikariPool-1 - Connection is not available, request timed out after 30000ms.
(A Java exception is an error object that is "thrown" and stops the current operation unless someone catches it - like a JS throw.)
- that is the only error pool exhaustion produces - and only after 30 s. Until then requests just wait. Hence: exhaustion presents as latency.
leak-detection-thresholdis the cheapest diagnostic you have: it prints the stack of whoever has held a connection too long - usually a method that holds a transaction open across a remote call.
The shape of an exhaustion incident
The most common cause is not "too much traffic". It is holding a connection while waiting on something else. In the Java below, @Transactional is an annotation (a label Spring reads) that means "run this method in one database transaction" - one all-or-nothing unit of work - so the connection is held from the first query to the end of the method:
@Transactional
public Receipt checkout(Cart c) {
orders.save(c); // takes a connection from the pool
payments.authorize(c); // HTTP call to another service - connection still held
return receipts.save(c); // commit, connection returned
}
If payments.authorize takes 8 s instead of 40 ms, every checkout holds a DB connection for 8 s. 20 connections / 8 s = 2.5 checkouts per second is the new capacity of the whole database pool - for every endpoint that needs the database, not just checkout. The list page, which never calls payments, now waits behind them. One slow dependency, and the whole service is slow.
Little's law: the arithmetic of pools
Little's law is a queueing rule: the number of things in a system equals the rate they arrive times how long each stays.
concurrency = throughput x latency (L = lambda x W)
20 connections, each held 8 s -> max 2.5 requests/s through the pool
20 connections, each held 40 ms -> max 500 requests/s
Use it both ways: to see what capacity a pool gives you, and to see what latency does to that capacity. It is also why fixing latency (a timeout) restores throughput when "adding connections" does not.
Bigger is usually worse
The intuitive fix - "raise maximum-pool-size from 20 to 200" - usually makes things slower:
- The database has finite parallelism: a few cores, a few disks. 200 active queries do not run 10x faster than 20; they contend for CPU, locks and buffer cache, and each takes longer.
- Every connection costs the database memory (a PostgreSQL backend process is ~5-10 MB) and context switches.
- Multiply by replicas: 10 pods x 200 connections = 2000 connections against a database with
max_connections=500. - Queueing at the pool is cheap and fair; queueing inside the database is not.
HikariCP's own guidance starts from connections = (cores x 2) + effective spindles - for a 4-core database, about 10. Measure, but start small. A pool of 10 routinely outperforms a pool of 100 on the same database.
The fix for exhaustion caused by slow downstream calls is not a bigger pool: it is a timeout on the call, and not holding the connection across it.
HTTP client pools
HTTP clients pool connections too, per route (host:port):
Apache HttpClient 5 (RestTemplate/RestClient with HttpComponents)
maxTotal 25, defaultMaxPerRoute 5 <- 5 concurrent calls to one service, by default!
connectionRequestTimeout 3 min waiting for a pooled connection
Reactor Netty (WebClient) pool max: 2 x CPUs, min 16; pendingAcquireTimeout 45 s
OkHttp maxIdle 5, but no per-route concurrency limit
An HTTP pool of 5 per route is exhaustion waiting to happen under load - threads queue for a connection with no error for up to three minutes. Size it for your concurrency, and set the pool-wait timeout short.
What to watch
hikaricp_connections_pending > 0 for more than a few seconds the pool is the bottleneck
hikaricp_connections_active == max saturated
hikaricp_connections_acquire_seconds_max how long borrowing takes
thread dump: many threads in HikariPool.getConnection the same thing, from inside
What you can now do
- Explain why an exhausted pool shows as latency, not errors, and read the HikariCP settings.
- Use Little's law to compute a pool's capacity from how long connections are held.
- Argue for a small pool plus a timeout instead of a bigger pool.