OnCallReady

Lesson 21.15 · Spring Boot Runtime, Resilience & Python Ops · 12 min read

Connection pools: why exhaustion is latency, not errors

In plain words

Imagine a library with ten copies of a popular book. Borrowing is quick while copies are on the shelf. If everybody who borrows a copy takes it home for a week instead of an hour, the eleventh person waits at the desk, and so does everyone after them. Nobody gets an error; they just wait longer and longer.

A database connection pool is that shelf. HikariCP keeps maximum-pool-size connections open and lends them to threads. When all are lent out, threads wait up to connection-timeout, 30 s by default, and only then get an exception. So exhaustion looks like latency, not errors. The usual cause is not too many readers but readers holding a copy while they wait for something else, like orders calling payments inside a transaction.

What a pool does

The problem. A service slows to 25-second responses with zero errors and idle CPU. More often than not a pool is full: all its connections are lent out and every new request waits in line. It looks like "the database is slow" and it is not.

What you need to know already: TCP handshakes and TLS (9.1, 9.15), threads and thread dumps (20.20, 20.22), the metrics page (21.11), latency percentiles (0.2).

Opening a database connection costs a TCP handshake, TLS, authentication and a server process or thread: tens of milliseconds. A pool keeps N connections open and lends them to threads. A thread that wants one when all N are lent out waits.

HikariCP settings that matter

spring.datasource.hikari.maximum-pool-size=10     N. Default 10.
spring.datasource.hikari.minimum-idle=10          default = maximum-pool-size (a fixed pool - recommended)
spring.datasource.hikari.connection-timeout=30000 ms a thread waits to BORROW before it gets an exception. Default 30 s.
spring.datasource.hikari.max-lifetime=1800000     retire connections before the DB or a firewall kills them. Default 30 min.
spring.datasource.hikari.idle-timeout=600000      only matters when minimum-idle < max
spring.datasource.hikari.leak-detection-threshold=20000   log a stack trace if a connection is held > 20 s
spring.datasource.hikari.keepalive-time=120000    ping idle connections (NAT / LB idle timeouts)

When a borrow times out:

java.sql.SQLTransientConnectionException: HikariPool-1 - Connection is not available, request timed out after 30000ms.

(A Java exception is an error object that is "thrown" and stops the current operation unless someone catches it - like a JS throw.)

The shape of an exhaustion incident

The most common cause is not "too much traffic". It is holding a connection while waiting on something else. In the Java below, @Transactional is an annotation (a label Spring reads) that means "run this method in one database transaction" - one all-or-nothing unit of work - so the connection is held from the first query to the end of the method:

@Transactional
public Receipt checkout(Cart c) {
    orders.save(c);                    // takes a connection from the pool
    payments.authorize(c);             // HTTP call to another service - connection still held
    return receipts.save(c);           // commit, connection returned
}

If payments.authorize takes 8 s instead of 40 ms, every checkout holds a DB connection for 8 s. 20 connections / 8 s = 2.5 checkouts per second is the new capacity of the whole database pool - for every endpoint that needs the database, not just checkout. The list page, which never calls payments, now waits behind them. One slow dependency, and the whole service is slow.

Little's law: the arithmetic of pools

Little's law is a queueing rule: the number of things in a system equals the rate they arrive times how long each stays.

concurrency = throughput x latency          (L = lambda x W)

20 connections, each held 8 s        ->   max 2.5 requests/s through the pool
20 connections, each held 40 ms      ->   max 500 requests/s

Use it both ways: to see what capacity a pool gives you, and to see what latency does to that capacity. It is also why fixing latency (a timeout) restores throughput when "adding connections" does not.

Bigger is usually worse

The intuitive fix - "raise maximum-pool-size from 20 to 200" - usually makes things slower:

HikariCP's own guidance starts from connections = (cores x 2) + effective spindles - for a 4-core database, about 10. Measure, but start small. A pool of 10 routinely outperforms a pool of 100 on the same database.

The fix for exhaustion caused by slow downstream calls is not a bigger pool: it is a timeout on the call, and not holding the connection across it.

HTTP client pools

HTTP clients pool connections too, per route (host:port):

Apache HttpClient 5 (RestTemplate/RestClient with HttpComponents)
  maxTotal 25, defaultMaxPerRoute 5            <- 5 concurrent calls to one service, by default!
  connectionRequestTimeout 3 min               waiting for a pooled connection
Reactor Netty (WebClient)     pool max: 2 x CPUs, min 16; pendingAcquireTimeout 45 s
OkHttp                        maxIdle 5, but no per-route concurrency limit

An HTTP pool of 5 per route is exhaustion waiting to happen under load - threads queue for a connection with no error for up to three minutes. Size it for your concurrency, and set the pool-wait timeout short.

What to watch

hikaricp_connections_pending > 0 for more than a few seconds    the pool is the bottleneck
hikaricp_connections_active == max                              saturated
hikaricp_connections_acquire_seconds_max                        how long borrowing takes
thread dump: many threads in HikariPool.getConnection           the same thing, from inside

What you can now do

Why it helps

When a service is slow with no errors and idle CPU, the pool is the first suspect, and hikaricp_connections_pending confirms it in seconds. The instinctive fix, raising the pool from 20 to 200, usually makes it worse: 10 pods times 200 connections is 2000 against a database with max_connections=500. Knowing Little's law lets you explain in a war room why a timeout on the slow call restores throughput and a bigger pool doesn't.

You'll also review these settings in Deployments and meet their limits in the cloud: managed-database connection limits, and NAT or load balancer idle timeouts that kill pooled connections, which is what max-lifetime and keepalive-time are for. HTTP client pools have the same problem with smaller defaults, 5 connections per route in Apache HttpClient.

FAQ

Why not just make the pool bigger?

Because the database has finite parallelism: a few cores and disks. 200 concurrent queries don't run faster than 20; they contend for CPU, locks and cache, and each gets slower. Each connection also costs database memory, around 5-10 MB per PostgreSQL backend, and multiplies by replicas. HikariCP's guidance starts at about cores times 2 plus spindles, around 10 for a 4-core database. If exhaustion comes from slow downstream calls, the fix is a timeout, not more connections.

What is Little's law and why does it matter here?

Concurrency equals throughput times latency. A pool of 20 connections, each held for 40 ms, can serve up to 500 requests per second. If each is held for 8 s because of a slow downstream call inside a transaction, the same pool serves only 2.5 per second, for every endpoint that needs the database. It shows why latency is the lever: cutting hold time with a timeout restores capacity that adding connections cannot.

What does leak-detection-threshold do?

It makes Hikari log a warning with a stack trace when a connection has been held longer than the threshold, say 20 seconds. The stack shows the code that borrowed it, which is usually a method holding a transaction open across a remote call or a connection not returned in an error path. It is cheap enough to enable in production and is often the fastest way to find the cause of exhaustion.

What are max-lifetime and keepalive-time for?

Connections that stay open for a long time get killed by things outside the app: the database's own limits, firewalls, NAT gateways and load balancers with idle timeouts. max-lifetime, 30 minutes by default, retires connections before those kill them. keepalive-time pings idle connections so an idle timeout, such as the 4-minute default some cloud load balancers and NAT gateways use, does not silently drop them. Symptoms of getting it wrong: occasional "connection reset" errors on the first query after idle.

Do HTTP clients have pools too?

Yes, per route, meaning per host and port. Apache HttpClient 5 defaults to 25 total and only 5 per route, and threads wait up to 3 minutes for a pooled connection by default. So a service calling payments with 50 concurrent requests has 45 threads queued with no error. Reactor Netty's pool is bigger. Size HTTP pools for your real concurrency to each dependency, and set a short pool-wait timeout.

In an interview Mid

A service's latency jumped to 30 seconds but the error rate is flat and CPU is idle. What do you suspect?

A full pool. A connection pool (HikariCP for the database) lends a fixed number of connections; when all are lent out, the next thread waits - up to connection-timeout, 30 s by default - and then usually gets one. So exhaustion shows as latency, not errors: the only error is Connection is not available, request timed out after 30000ms, and only after 30 s.

Confirm: hikaricp_connections_active = max and hikaricp_connections_pending > 0; a thread dump full of threads in HikariPool.getConnection.

The usual cause is holding a connection while waiting on something else - a @Transactional method that calls another service over HTTP. Little's law: 20 connections held 8 s each = 2.5 requests/s for the whole pool, for every endpoint that needs the database.

The fix is a timeout on the slow call and not holding the connection across it - not a bigger pool: the database has finite parallelism, and 10 pods x 200 connections can exceed max_connections. leak-detection-threshold shows who holds connections too long.

Also asked: What is a connection pool and why do applications use one? · Why is a bigger connection pool often slower? · How would you size a database connection pool for a service with several replicas?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.