Retries
The problem. Timeouts turn hangs into errors - now what? Retry, stop calling, or limit the damage? Done naively, each of these makes outages worse. This lesson is the small toolkit of resilience patterns and the rules that keep them safe.
What you need to know already: timeouts and budgets (21.18), pools and Little's law (21.15), HTTP status codes (9.21), thread pools (21.3).
A retry is a bet that the failure was transient (temporary). Rules:
- Only retry idempotent operations (idempotent = doing it twice has the same effect as doing it once). A GET, a PUT of a full resource, a DELETE: doing it twice is the same as once. A POST that charges a card is not - unless the API takes an idempotency key (the client generates a unique id per logical operation; the server remembers it and returns the first result for duplicates). Payment providers all support this; use it.
- Only retry retryable failures: connect errors, timeouts, 502/503/504, 429 (honour
Retry-After). Not 400, 401, 404, 409. - Back off exponentially: wait base, 2 x base, 4 x base... with a cap.
- Add jitter (randomness in the wait). Randomise each wait (full jitter:
sleep = random(0, backoff)). - Bound the total. A retry budget or a maximum number of attempts, inside the caller's deadline.
Resilience4j is the Java library Spring apps use for these patterns; each one is configured with properties per named target ("instance", here payments):
resilience4j.retry.instances.payments.max-attempts=3
resilience4j.retry.instances.payments.wait-duration=200ms
resilience4j.retry.instances.payments.enable-exponential-backoff=true
resilience4j.retry.instances.payments.exponential-backoff-multiplier=2
resilience4j.retry.instances.payments.enable-randomized-wait=true
Retry storms
A dependency has a 10-second outage. 600 clients are retrying every second on the dot. When it comes back it receives all 600 at once, plus normal traffic - several times its capacity - a retry storm. Overloaded, it times out most of them, which fail and retry - in sync again - one second later. The service never gets below its capacity long enough to recover. The retries are the outage now.
It gets worse with layers: if the edge retries 3 times, the API retries 3 times and the service behind it retries 3 times, one user request can become 27 requests at the bottom. Retry at one layer, usually the one closest to the failure.
Jitter breaks the synchronisation: with random waits, the retry load arrives as a smear instead of a wave, stays under capacity, and the service recovers. This chapter's stormsim (a simulator tool) lets you watch both.
Circuit breakers
A circuit breaker (named after the electrical fuse box switch) stops you calling something that is failing, so you fail fast (and free the thread) instead of waiting on it:
CLOSED normal. Calls go through; outcomes are recorded in a sliding window.
failure rate (or slow-call rate) >= threshold over >= minimum calls -> OPEN
OPEN every call is rejected immediately (CallNotPermittedException) - a fallback or a fast 503.
after wait-duration-in-open-state -> HALF_OPEN
HALF_OPEN a few trial calls are let through.
they succeed -> CLOSED; they fail -> OPEN again
resilience4j.circuitbreaker.instances.payments.sliding-window-size=20
resilience4j.circuitbreaker.instances.payments.minimum-number-of-calls=10
resilience4j.circuitbreaker.instances.payments.failure-rate-threshold=50
resilience4j.circuitbreaker.instances.payments.slow-call-duration-threshold=2s
resilience4j.circuitbreaker.instances.payments.wait-duration-in-open-state=10s
resilience4j.circuitbreaker.instances.payments.permitted-number-of-calls-in-half-open-state=3
Defaults to know: window 100 calls, minimum 100 calls, failure threshold 50%, slow-call threshold 60 s, open for 60 s. The defaults are slow to trip on a low-traffic service - tune them.
(A fallback = the answer you give instead when the real call is not made - a cached value, a "try later" message.)
A circuit breaker needs a timeout underneath it. It records outcomes when a call completes. A call to a hung service with no read timeout never completes, so it is never counted, and the breaker never opens. Timeout first, then breaker.
Observability:
$ curl -s localhost:8080/actuator/circuitbreakers | jq .
{
"circuitBreakers": {
"payments": {
"failureRate": "100.0%",
...
"notPermittedCalls": 180,
"state": "OPEN"
...
# the same state as a metric, and in the log:
resilience4j_circuitbreaker_state{name="payments",state="open"} 1.0
CircuitBreaker 'payments' changed state from CLOSED to OPEN
Bulkheads
A ship's bulkheads (the watertight walls between compartments) stop one flooded compartment sinking the ship. In a service: separate, bounded resources per dependency, so one slow downstream cannot consume every thread. A dedicated thread pool (or a semaphore - a counter of allowed concurrent calls) of 20 for payments calls means at most 20 threads can be stuck on payments; the other 180 Tomcat threads keep serving everything else.
resilience4j.bulkhead.instances.payments.max-concurrent-calls=20
resilience4j.bulkhead.instances.payments.max-wait-duration=0 # reject immediately when full
Backpressure
Backpressure = pushing back on callers when you are full. When a service is saturated it must say so instead of queueing silently until it dies: reject with 429/503 (and Retry-After), bound its queues, and let the caller slow down. An unbounded queue in front of a slow consumer is not buffering; it is a memory leak with a delay and ever-growing latency.
unbounded queue latency grows without limit, then OOM, nothing ever fails fast
bounded queue + reject the excess fails fast and visibly; the rest is served at normal latency
Putting it together for a call to payments
bulkhead (max 20 concurrent) -> circuit breaker -> retry (2 attempts, backoff+jitter, idempotency key)
-> timeout (read 2 s, total 3 s) -> HTTP client pool (sized for 20, short pool-wait timeout)
fallback: "payment pending" + asynchronous reconciliation, or a fast 503
What you can now do
- State the retry rules: idempotent only, retryable failures only, backoff + jitter, bounded.
- Explain a retry storm and why jitter and retrying at one layer stop it.
- Describe circuit breaker states, bulkheads and backpressure, and why a breaker needs a timeout under it.