OnCallReady

Lesson 21.22 · Spring Boot Runtime, Resilience & Python Ops · 13 min read

Retries, jitter, circuit breakers, bulkheads, backpressure

In plain words

Imagine a crowd outside a shop that just closed for ten minutes. If everyone retries the door every second exactly, the moment it opens they all rush in together and the shop is overwhelmed again. If each person waits a random amount of time, they arrive a few at a time and the shop copes.

That is retries with jitter. A circuit breaker is the shop putting up a "back in 10 minutes" sign so people stop trying for a while. A bulkhead is having separate queues for different counters, so a slow counter can't block the whole shop. Backpressure is a bouncer saying "full, come back later" instead of letting the queue grow forever. Resilience4j gives a Spring service all four around its call to payments.

Retries

The problem. Timeouts turn hangs into errors - now what? Retry, stop calling, or limit the damage? Done naively, each of these makes outages worse. This lesson is the small toolkit of resilience patterns and the rules that keep them safe.

What you need to know already: timeouts and budgets (21.18), pools and Little's law (21.15), HTTP status codes (9.21), thread pools (21.3).

A retry is a bet that the failure was transient (temporary). Rules:

  1. Only retry idempotent operations (idempotent = doing it twice has the same effect as doing it once). A GET, a PUT of a full resource, a DELETE: doing it twice is the same as once. A POST that charges a card is not - unless the API takes an idempotency key (the client generates a unique id per logical operation; the server remembers it and returns the first result for duplicates). Payment providers all support this; use it.
  2. Only retry retryable failures: connect errors, timeouts, 502/503/504, 429 (honour Retry-After). Not 400, 401, 404, 409.
  3. Back off exponentially: wait base, 2 x base, 4 x base... with a cap.
  4. Add jitter (randomness in the wait). Randomise each wait (full jitter: sleep = random(0, backoff)).
  5. Bound the total. A retry budget or a maximum number of attempts, inside the caller's deadline.

Resilience4j is the Java library Spring apps use for these patterns; each one is configured with properties per named target ("instance", here payments):

resilience4j.retry.instances.payments.max-attempts=3
resilience4j.retry.instances.payments.wait-duration=200ms
resilience4j.retry.instances.payments.enable-exponential-backoff=true
resilience4j.retry.instances.payments.exponential-backoff-multiplier=2
resilience4j.retry.instances.payments.enable-randomized-wait=true

Retry storms

A dependency has a 10-second outage. 600 clients are retrying every second on the dot. When it comes back it receives all 600 at once, plus normal traffic - several times its capacity - a retry storm. Overloaded, it times out most of them, which fail and retry - in sync again - one second later. The service never gets below its capacity long enough to recover. The retries are the outage now.

It gets worse with layers: if the edge retries 3 times, the API retries 3 times and the service behind it retries 3 times, one user request can become 27 requests at the bottom. Retry at one layer, usually the one closest to the failure.

Jitter breaks the synchronisation: with random waits, the retry load arrives as a smear instead of a wave, stays under capacity, and the service recovers. This chapter's stormsim (a simulator tool) lets you watch both.

Circuit breakers

A circuit breaker (named after the electrical fuse box switch) stops you calling something that is failing, so you fail fast (and free the thread) instead of waiting on it:

CLOSED      normal. Calls go through; outcomes are recorded in a sliding window.
            failure rate (or slow-call rate) >= threshold over >= minimum calls  -> OPEN
OPEN        every call is rejected immediately (CallNotPermittedException) - a fallback or a fast 503.
            after wait-duration-in-open-state                                    -> HALF_OPEN
HALF_OPEN   a few trial calls are let through.
            they succeed -> CLOSED;  they fail -> OPEN again
resilience4j.circuitbreaker.instances.payments.sliding-window-size=20
resilience4j.circuitbreaker.instances.payments.minimum-number-of-calls=10
resilience4j.circuitbreaker.instances.payments.failure-rate-threshold=50
resilience4j.circuitbreaker.instances.payments.slow-call-duration-threshold=2s
resilience4j.circuitbreaker.instances.payments.wait-duration-in-open-state=10s
resilience4j.circuitbreaker.instances.payments.permitted-number-of-calls-in-half-open-state=3

Defaults to know: window 100 calls, minimum 100 calls, failure threshold 50%, slow-call threshold 60 s, open for 60 s. The defaults are slow to trip on a low-traffic service - tune them.

(A fallback = the answer you give instead when the real call is not made - a cached value, a "try later" message.)

A circuit breaker needs a timeout underneath it. It records outcomes when a call completes. A call to a hung service with no read timeout never completes, so it is never counted, and the breaker never opens. Timeout first, then breaker.

Observability:

$ curl -s localhost:8080/actuator/circuitbreakers | jq .
{
  "circuitBreakers": {
    "payments": {
      "failureRate": "100.0%",
      ...
      "notPermittedCalls": 180,
      "state": "OPEN"
...
# the same state as a metric, and in the log:
resilience4j_circuitbreaker_state{name="payments",state="open"} 1.0
CircuitBreaker 'payments' changed state from CLOSED to OPEN

Bulkheads

A ship's bulkheads (the watertight walls between compartments) stop one flooded compartment sinking the ship. In a service: separate, bounded resources per dependency, so one slow downstream cannot consume every thread. A dedicated thread pool (or a semaphore - a counter of allowed concurrent calls) of 20 for payments calls means at most 20 threads can be stuck on payments; the other 180 Tomcat threads keep serving everything else.

resilience4j.bulkhead.instances.payments.max-concurrent-calls=20
resilience4j.bulkhead.instances.payments.max-wait-duration=0      # reject immediately when full

Backpressure

Backpressure = pushing back on callers when you are full. When a service is saturated it must say so instead of queueing silently until it dies: reject with 429/503 (and Retry-After), bound its queues, and let the caller slow down. An unbounded queue in front of a slow consumer is not buffering; it is a memory leak with a delay and ever-growing latency.

unbounded queue      latency grows without limit, then OOM, nothing ever fails fast
bounded queue + reject    the excess fails fast and visibly; the rest is served at normal latency

Putting it together for a call to payments

bulkhead (max 20 concurrent)  ->  circuit breaker  ->  retry (2 attempts, backoff+jitter, idempotency key)
  ->  timeout (read 2 s, total 3 s)  ->  HTTP client pool (sized for 20, short pool-wait timeout)
fallback: "payment pending" + asynchronous reconciliation, or a fast 503

What you can now do

Why it helps

Retry storms, where the retries themselves become the outage, happen in real platforms after every significant dependency blip, and they are why a 10-second outage can turn into an hour. Knowing that retries need backoff, jitter, idempotency and a single retry layer lets you spot the risk in a design review, including in things you configure yourself: ingress retries, retries built into cloud SDKs, and your own scripts.

Circuit breakers and bulkheads also appear in incident timelines you'll read: "CircuitBreaker 'payments' changed state from CLOSED to OPEN" in the logs, notPermittedCalls in /actuator/circuitbreakers. You'll know that a breaker needs a timeout underneath to work at all, and that Resilience4j's defaults barely trip on a low-traffic service. These are also classic SRE interview questions.

Commands in this lesson

curl

FAQ

Which operations are safe to retry?

Idempotent ones, where doing it twice has the same effect as once: GET, HEAD, PUT of a full resource, DELETE. A POST that charges a card or creates an order is not, unless the API supports an idempotency key: the client sends a unique ID per logical operation and the server returns the first result for duplicates. Also only retry retryable failures: connect errors, timeouts, 502, 503, 504, and 429 honouring Retry-After, not 400, 401, 404 or 409.

Why add jitter to backoff?

Without jitter, clients that failed at the same moment retry at the same moments: 1 s, 2 s, 4 s later, all together. The recovering service gets waves of synchronised load above its capacity, fails them, and the waves repeat. Jitter randomises each wait, for example full jitter with sleep = random(0, backoff), so retries arrive as a smear under capacity and the service can recover. Exponential backoff alone is not enough.

Why does a circuit breaker need a timeout?

A breaker counts outcomes when calls complete. A call to a hung service with no read timeout never completes, so it is never recorded as a failure or slow call, and the breaker never opens, while threads pile up. With a read timeout underneath, each hung call becomes a timeout exception after a bounded time, the failure rate crosses the threshold, and the breaker opens. So the order is timeout first, then breaker.

Why is retrying at several layers dangerous?

Retries multiply. If the edge retries 3 times, the API gateway 3 times and the service 3 times, one user request can become 27 requests at the bottom, exactly when the bottom is struggling. Retry at one layer, usually the one closest to the failure, and make the others pass errors through. Check the defaults of everything in the path: ingress controllers, service meshes and SDKs often retry on their own.

What is backpressure in practice?

A saturated service saying so instead of queueing silently. Concretely: bounded queues, rejecting excess work quickly with 429 or 503 and Retry-After, and callers slowing down in response. An unbounded queue in front of a slow consumer isn't buffering; latency grows without limit and eventually memory runs out, and nothing fails fast. With a bound, the excess fails visibly and the rest is served at normal latency.

In an interview Mid

What is a circuit breaker and how does it work?

A circuit breaker stops calling a dependency that is failing, so you fail fast (and free the thread) instead of waiting on it:

In Resilience4j it is properties per instance (sliding-window-size, minimum-number-of-calls, failure-rate-threshold). The defaults (100 calls, 60 s slow-call threshold) are slow to trip on a low-traffic service - tune them.

The catch: it needs a timeout underneath. It counts a call when the call completes; a hung call with no read timeout never completes, so the breaker never opens. Timeout first, then breaker, plus a bulkhead so only a bounded number of threads can wait on that dependency.

Also asked: When is it safe to retry a failed request? · Explain retry storms and how jitter prevents them. · What is the bulkhead pattern?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.