Three different timeouts
The problem. One downstream service hangs, and your service - which was healthy - stops answering too, then the replicas fall over one by one. The missing piece is almost always a timeout that was never set.
What you need to know already: TCP connect and the three ways a connection fails (9.1), curl -w timings (9.27), pools (21.15), SLOs (0.8), strace on orders' read() (3.12).
A timeout = the longest you are willing to wait before giving up with an error. There are several, for different phases of a call:
connect timeout establishing the TCP (and TLS) connection "is anyone there?"
read / socket max silence while waiting for bytes, once connected "is it still answering?"
total / request the whole operation: connect + send + wait + read + retries "is it done yet?"
(pool wait) waiting for a free pooled connection before any of the above
They fail differently. A connect timeout means the host is not answering at all (down, dropped by a firewall) - usually fast to diagnose. A read timeout means it accepted the connection and then went quiet: the far end is overloaded, blocked on its own dependency, or deadlocked. The read timeout is the one that protects your threads.
The defaults: often none
The most important fact in this block: many clients default to no read timeout at all. A thread that calls a hung service waits forever.
java.net.HttpURLConnection (RestTemplate default factory) connect: infinite read: infinite
java.net.http.HttpClient (JDK 11+) connect: none request: none
Apache HttpClient 5 connect: 3 min response: none
Spring WebClient / Reactor Netty connect: 30 s response: none
OkHttp connect/read/write: 10 s (the exception)
PostgreSQL JDBC socketTimeout 0 = infinite
python requests none, unless you pass timeout=
curl connect: 300 s total: none (use -m)
Read that table as: unless someone set it, assume there is no timeout.
On this box: /etc/orders/app.conf has had http.client.read.timeout.ms=0 since Ch 3 (3.12), where strace showed orders blocked in read(). Zero means infinite. This chapter is where that line becomes an incident.
What "no timeout" does to a service
With no read timeout, a hung dependency turns each request into a thread that never comes back. Tomcat's 200 threads fill up in 200 requests; after that new requests queue in the accept backlog (connections the kernel accepted but the app has not picked up yet, 9.8) and eventually time out at the client or the load balancer. The service's own health check - if it runs on the same thread pool, or needs the same exhausted DB pool - starts failing. The orchestrator restarts it; its traffic moves to the other replicas; they fill up the same way. A cascading failure (one failure causing the next, and so on) that started as one slow dependency, and the root cause is a missing number in a config file.
A timeout converts "hang" into "error": the thread comes back after N ms with an exception you can count, alert on, retry, or answer with a fallback. That is strictly better.
Setting them
# Spring Boot 3.4+ (RestClient / RestTemplate built by Boot)
spring.http.client.connect-timeout=2s
spring.http.client.read-timeout=2s
# RestTemplateBuilder
builder.connectTimeout(Duration.ofSeconds(2)).readTimeout(Duration.ofSeconds(2))
# JDK HttpClient
HttpClient.newBuilder().connectTimeout(Duration.ofSeconds(2))
HttpRequest.newBuilder(uri).timeout(Duration.ofSeconds(2)) # per request
# Resilience4j (the Java resilience library, lesson 21.22) TimeLimiter - a total time bound around async calls
resilience4j.timelimiter.instances.payments.timeout-duration=2s
# JDBC
spring.datasource.hikari.data-source-properties.socketTimeout=10 # PostgreSQL, seconds
The budget
An API with an SLO of 2 s that calls three services in sequence cannot give each of them 2 s. The budget has to be divided, with room left for your own work and a retry:
SLO 2000 ms
own work ~150 ms (parsing, DB, serialisation)
A (auth) p99 40 ms -> timeout 200 ms
B (pricing) p99 150 ms -> timeout 500 ms
C (inventory) p99 300 ms -> timeout 800 ms
---------
worst case 150 + 200 + 500 + 800 = 1650 ms < 2000 - leaves ~350 ms, no room for a retry of C
Principles:
- A timeout is set from the dependency's latency distribution (a few times its p99), not from your own SLO.
- The sum of timeouts on the sequential path (times attempts, if you retry) must fit inside your SLO - otherwise your timeouts cannot protect it.
- Propagate the deadline. If A already used 1.2 s of a 2 s request, the call to B should get at most the 0.8 s that is left, not its full 500 ms default. gRPC (17.20) deadlines do this for you; with HTTP you pass a header or a context and compute the remaining time.
- Upstream timeouts must be longer than downstream ones: nginx's
proxy_read_timeout(default 60 s) > the app's total > each client's read timeout. Otherwise the outer layer gives up while inner layers keep working on a request nobody will read.
Measuring from the outside: hey
hey is a small load generator: it sends many HTTP requests at once and prints a latency summary.
$ hey -n 2000 -c 50 http://localhost:8080/api/orders
Summary:
Total: 0.4014 secs
Slowest: 0.0119 secs
Fastest: 0.0032 secs
Average: 0.0099 secs
Requests/sec: 4982.5350
Response time histogram:
0.003 [1] |
...
0.010 [865] |■■■■■■■■■■■■■■■■■■■■■■■■■■■■■■■■■■■■■■■■
Latency distribution:
50% in 0.0099 secs
90% in 0.0108 secs
99% in 0.0114 secs
Status code distribution:
[200] 2000 responses
-n requests, -c concurrent workers, -z 30s for a duration, -m POST, -t per-request timeout (default 20 s - requests slower than that show up in "Error distribution" as context deadline exceeded, and keep running on the server). hey is closed-loop: each worker waits for its response before sending the next, so when the server slows down, hey slows down with it. It measures latency under a concurrency level; it is not a model of a real crowd.
What you can now do
- Tell connect, read, total and pool-wait timeouts apart, and assume "no timeout" unless set.
- Split an SLO into per-call timeouts and propagate the remaining deadline.
- Load-test with hey and read its latency distribution.