OnCallReady

Lesson 21.18 · Spring Boot Runtime, Resilience & Python Ops · 12 min read

Timeouts: connect, read, total - and the budget

In plain words

Imagine phoning a friend. First you wait for them to pick up: if nobody answers after a while, you hang up. Once they answer, you wait for them to reply when you ask something: if they go completely silent for a minute, you hang up. And your mum said the whole call can't be longer than ten minutes, however it's going.

Those are the three timeouts. Connect timeout is "does anyone pick up?", read timeout is "has it gone silent?", and the total timeout bounds the whole call including retries. The dangerous part: many HTTP clients default to no read timeout, so the call waits forever. http.client.read.timeout.ms=0 in /etc/orders/app.conf is exactly that, and it is why orders was blocked in read() back in Ch 3 (3.12).

Three different timeouts

The problem. One downstream service hangs, and your service - which was healthy - stops answering too, then the replicas fall over one by one. The missing piece is almost always a timeout that was never set.

What you need to know already: TCP connect and the three ways a connection fails (9.1), curl -w timings (9.27), pools (21.15), SLOs (0.8), strace on orders' read() (3.12).

A timeout = the longest you are willing to wait before giving up with an error. There are several, for different phases of a call:

connect timeout     establishing the TCP (and TLS) connection     "is anyone there?"
read / socket       max silence while waiting for bytes, once connected   "is it still answering?"
total / request     the whole operation: connect + send + wait + read + retries   "is it done yet?"
(pool wait)         waiting for a free pooled connection before any of the above

They fail differently. A connect timeout means the host is not answering at all (down, dropped by a firewall) - usually fast to diagnose. A read timeout means it accepted the connection and then went quiet: the far end is overloaded, blocked on its own dependency, or deadlocked. The read timeout is the one that protects your threads.

The defaults: often none

The most important fact in this block: many clients default to no read timeout at all. A thread that calls a hung service waits forever.

java.net.HttpURLConnection (RestTemplate default factory)   connect: infinite   read: infinite
java.net.http.HttpClient (JDK 11+)                           connect: none       request: none
Apache HttpClient 5                                          connect: 3 min      response: none
Spring WebClient / Reactor Netty                             connect: 30 s       response: none
OkHttp                                                        connect/read/write: 10 s  (the exception)
PostgreSQL JDBC socketTimeout                                 0 = infinite
python requests                                               none, unless you pass timeout=
curl                                                          connect: 300 s   total: none (use -m)

Read that table as: unless someone set it, assume there is no timeout.

On this box: /etc/orders/app.conf has had http.client.read.timeout.ms=0 since Ch 3 (3.12), where strace showed orders blocked in read(). Zero means infinite. This chapter is where that line becomes an incident.

What "no timeout" does to a service

With no read timeout, a hung dependency turns each request into a thread that never comes back. Tomcat's 200 threads fill up in 200 requests; after that new requests queue in the accept backlog (connections the kernel accepted but the app has not picked up yet, 9.8) and eventually time out at the client or the load balancer. The service's own health check - if it runs on the same thread pool, or needs the same exhausted DB pool - starts failing. The orchestrator restarts it; its traffic moves to the other replicas; they fill up the same way. A cascading failure (one failure causing the next, and so on) that started as one slow dependency, and the root cause is a missing number in a config file.

A timeout converts "hang" into "error": the thread comes back after N ms with an exception you can count, alert on, retry, or answer with a fallback. That is strictly better.

Setting them

# Spring Boot 3.4+ (RestClient / RestTemplate built by Boot)
spring.http.client.connect-timeout=2s
spring.http.client.read-timeout=2s

# RestTemplateBuilder
builder.connectTimeout(Duration.ofSeconds(2)).readTimeout(Duration.ofSeconds(2))

# JDK HttpClient
HttpClient.newBuilder().connectTimeout(Duration.ofSeconds(2))
HttpRequest.newBuilder(uri).timeout(Duration.ofSeconds(2))   # per request

# Resilience4j (the Java resilience library, lesson 21.22) TimeLimiter - a total time bound around async calls
resilience4j.timelimiter.instances.payments.timeout-duration=2s

# JDBC
spring.datasource.hikari.data-source-properties.socketTimeout=10    # PostgreSQL, seconds

The budget

An API with an SLO of 2 s that calls three services in sequence cannot give each of them 2 s. The budget has to be divided, with room left for your own work and a retry:

SLO                2000 ms
own work            ~150 ms   (parsing, DB, serialisation)
A (auth)            p99 40 ms    -> timeout 200 ms
B (pricing)         p99 150 ms   -> timeout 500 ms
C (inventory)       p99 300 ms   -> timeout 800 ms
                    ---------
worst case          150 + 200 + 500 + 800 = 1650 ms  < 2000 - leaves ~350 ms, no room for a retry of C

Principles:

Measuring from the outside: hey

hey is a small load generator: it sends many HTTP requests at once and prints a latency summary.

$ hey -n 2000 -c 50 http://localhost:8080/api/orders

Summary:
  Total:	0.4014 secs
  Slowest:	0.0119 secs
  Fastest:	0.0032 secs
  Average:	0.0099 secs
  Requests/sec:	4982.5350

Response time histogram:
  0.003 [1]	|
  ...
  0.010 [865]	|■■■■■■■■■■■■■■■■■■■■■■■■■■■■■■■■■■■■■■■■

Latency distribution:
  50% in 0.0099 secs
  90% in 0.0108 secs
  99% in 0.0114 secs

Status code distribution:
  [200]	2000 responses

-n requests, -c concurrent workers, -z 30s for a duration, -m POST, -t per-request timeout (default 20 s - requests slower than that show up in "Error distribution" as context deadline exceeded, and keep running on the server). hey is closed-loop: each worker waits for its response before sending the next, so when the server slows down, hey slows down with it. It measures latency under a concurrency level; it is not a model of a real crowd.

What you can now do

Why it helps

A missing read timeout is the root cause behind a large share of cascading outages: one hung dependency, threads that never come back, a pool that fills, and a service that stays "up" doing nothing. Checking timeouts is one of the highest-value things you can do in a design review, and "what's your read timeout to X?" is a question you should ask every team.

The budget part matters for SLOs you will own. If an API promises 2 seconds and calls three services in sequence, each with a 2-second timeout, the timeouts cannot protect the SLO. You'll also recognise layered timeouts: nginx or Application Gateway giving up at 60 s while the app keeps working on a request nobody will read, and how to set each layer.

Commands in this lesson

hey

FAQ

What is the difference between a read timeout and a total timeout?

A read timeout limits silence: how long the client waits for the next bytes once connected. A server that trickles one byte every few seconds never triggers it. A total or request timeout limits the whole operation: connect, send, waiting, reading the whole body, and any retries. You usually want both: a read timeout to catch hung servers quickly, and a total to guarantee an upper bound. Python's requests timeout is per socket read, not total.

Which clients have no read timeout by default?

Most of them. HttpURLConnection (RestTemplate's default), the JDK HttpClient, Apache HttpClient 5's response timeout, Reactor Netty's response timeout, the PostgreSQL JDBC socketTimeout and Python requests all default to none or infinite. OkHttp is the exception with 10 s. curl has a 300 s connect timeout and no total unless you pass -m. Assume there is no timeout unless someone set one.

How do I choose a timeout value?

From the dependency's latency distribution, not from your own SLO. A few times the dependency's p99 is a common starting point: if payments normally answers in 150 ms at p99, 500 ms is reasonable. Then check the budget: the sum of timeouts along a sequential path, multiplied by attempts if you retry, must fit inside your own SLO with room for your own work. If it doesn't, something must give: fewer retries, parallel calls, or a fallback.

Why should upstream timeouts be longer than downstream ones?

If the outer layer gives up first, the inner layers keep working on a request nobody will read, which wastes capacity during exactly the incident where it's scarce. So the proxy's timeout, like nginx's proxy_read_timeout (60 s default) or an Application Gateway request timeout, should be longer than the app's total, which should be longer than each client's read timeout. Better still, propagate the remaining deadline so each hop knows how much time is left.

What does hey measure, and what doesn't it?

hey -n 2000 -c 50 URL sends 2000 requests using 50 concurrent workers and reports latency percentiles and status codes. It is closed-loop: each worker waits for its response before sending the next, so when the server slows down, hey sends less. That measures latency at a fixed concurrency, not what a real crowd of independent users does, where arrivals keep coming. Its per-request timeout defaults to 20 s.

In an interview Mid

What timeouts should an HTTP client have, and how do you choose them?

Assume there is no timeout unless someone set it: HttpURLConnection and Reactor Netty have no read/response timeout by default, and on oncall-lab orders had http.client.read.timeout.ms=0 (infinite).

Choosing values - the budget: set each from the dependency's latency (a few times its p99), make the sum on the sequential path (times retries) fit inside your own SLO, propagate the deadline (pass on only what is left), and keep upstream timeouts longer than downstream ones - nginx proxy_read_timeout > the app's total > each client's read timeout.

Also asked: Explain timeout budgets across a chain of service calls. · A dependency hung and your service went down with it. Explain the mechanism. · What does a connect timeout tell you that a read timeout does not?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.