OnCallReady

Lesson 34.35 · Kubernetes: Ingress, Gateway API & Service Mesh · 13 min read

Traffic policy: VirtualService, DestinationRule, retries, timeouts, faults

In plain words

Think of a dispatcher at a taxi company. One sheet of rules says where each call goes: business customers to the premium fleet, everyone else split 80/20 between the old and new cars, and give up if no car answers within two seconds. A second sheet says how to treat each fleet: no more than ten jobs at once, and a car that fails three jobs in a row is sent to the garage for a while.

In Istio, the VirtualService is the first sheet: matches, weights, retries, timeouts and faults, where the first match wins. The DestinationRule is the second: subsets by label, connection pools (circuit breaking) and Outlier detection. Istio also retries some failures by default (2 attempts on 503-like errors), which is why the lesson ends with the retry storm: retries at every layer multiply the load on a struggling service.

Two objects, two moments

Istio's traffic API has two halves, applied at two moments of a request:

VirtualService     ROUTING: for requests to this host, which destination?
                   (matches, weights, retries, timeouts, faults, rewrites, mirrors)
DestinationRule    POLICY after routing: how to talk to that destination
                   (subsets = named groups of pods, load balancing, TLS,
                    connection pools, outlier detection)

What you need to know already: the Istio and mTLS lessons, Deployments with labels (15.6), canary releases with weights (16.27), HTTP 5xx codes and timeouts (9.22-9.23), and the edge timeouts of this chapter.

apiVersion: networking.istio.io/v1
kind: DestinationRule
metadata: {name: reviews, namespace: shop}
spec:
  host: reviews                       # the Service (short name = this namespace)
  subsets:
  - name: v1
    labels: {version: v1}             # pods of the Service with this label
  - name: v2
    labels: {version: v2}
---
apiVersion: networking.istio.io/v1
kind: VirtualService
metadata: {name: reviews, namespace: shop}
spec:
  hosts: [reviews]
  http:
  - match:
    - headers:
        x-canary: {exact: "true"}
    route:
    - destination: {host: reviews, subset: v2}
  - route:
    - destination: {host: reviews, subset: v1}
      weight: 90
    - destination: {host: reviews, subset: v2}
      weight: 10

The http list is ordered: the first match wins (unlike Gateway API's "most specific wins"). Put specific rules first, the catch-all last. A VS that matches nothing for a request gives a 404 with flag NR.

A route to a subset no DestinationRule defines is a 503 (flag NC, cluster not found) - istioctl analyze reports it as IST0101 before it hurts. A subset whose labels match no pod is a 503 no healthy upstream (flag UH).

Retries: there is a default

Even without any VirtualService, Istio retries: 2 attempts on connect-failure, refused-stream, unavailable, cancelled, retriable-status-codes - and the retriable status code is 503. So an upstream answering 503 is quietly called up to 3 times. Good for a flaky pod, bad when a dependency is overloaded.

http:
- route: [{destination: {host: inventory}}]
  timeout: 3s                 # the whole call, all attempts included
  retries:
    attempts: 3
    perTryTimeout: 1s         # each attempt
    retryOn: 5xx,gateway-error,connect-failure,reset

Rules of thumb:

Fault injection

The client sidecar can break things on purpose, to test the retries, timeouts and fallbacks you just wrote:

http:
- fault:
    delay: {percentage: {value: 50}, fixedDelay: 2s}       # flag DI
    abort: {percentage: {value: 10}, httpStatus: 503}      # flag FI, body "fault filter abort"
  route: [{destination: {host: ratings}}]

Faults happen in the client's sidecar before routing, so only callers in the mesh see them, and an injected delay does not count against that route's own timeout - the caller's timeout is what you are testing.

Outlier detection: take a bad pod out

One replica of five returns 500s (a bad node, a stuck cache). Load balancing keeps sending it 20% of traffic. Outlier detection ejects it:

spec:
  host: inventory
  trafficPolicy:
    outlierDetection:
      consecutive5xxErrors: 3      # 3 errors in a row from that host
      interval: 10s                # how often hosts are evaluated
      baseEjectionTime: 30s        # first ejection 30 s, then 60 s, 90 s...
      maxEjectionPercent: 50       # never eject more than half (default 10%, at least one host)

Each client proxy decides on its own (there is no global view): in istioctl proxy-config endpoints the ejected host shows OUTLIER CHECK FAILED, and the stat outlier_detection.ejections_active counts it.

Circuit breaking: limit what you send

connectionPool caps how much one client proxy sends to a destination:

trafficPolicy:
  connectionPool:
    tcp: {maxConnections: 1}
    http: {http1MaxPendingRequests: 1, maxRequestsPerConnection: 1}

Requests over the limit fail immediately with 503 and flag UO (upstream overflow) instead of piling up on a struggling service. It protects the server and frees the client's threads; the client must handle the fast 503 (a fallback, a cached answer). Load test it - fortio load -c 3 -qps 0 -n 30 URL from a pod in the mesh shows the overflow as Code 503.

In an interview: "VirtualService decides where a request goes - matches, weights, retries, timeouts, faults; DestinationRule decides how to talk to the destination - subsets, load balancing, connection pools, outlier detection. Istio retries 503s twice by default, which multiplies load across layers if every hop retries."

What you can now do:

Why it helps

Canaries, timeouts and retries are where a mesh pays for itself, and also where it causes incidents. A subset with no matching pods gives 503 no healthy upstream; a VirtualService that matches nothing gives 404 with flag NR; three layers of retries turn one failing service into a storm.

Knowing which object controls which behaviour lets you fix the right one, and the retry rule ("retry at one layer, and keep timeouts shorter as you go deeper") is one of the most useful reliability lessons there is, mesh or not. Fault injection lets you test how callers behave when a dependency is slow before it happens for real.

Commands in this lesson

kubectl

FAQ

Why do I get 503 with flag NC after adding a subset?

The VirtualService routes to a subset that no DestinationRule defines, so the proxy has no cluster for it. Create the DestinationRule with that subset (or fix the name). istioctl analyze reports it as IST0101 before you even send traffic.

Does Istio retry without me configuring anything?

Yes. The default policy retries twice on connection failures, refused streams, unavailable and reset-related errors, including 503. That hides blips, but it also means every layer in a call chain may retry, so set retries.attempts: 0 on the inner layers or tune it on purpose.

What is the difference between outlier detection and circuit breaking?

Outlier detection takes one bad endpoint out of the pool for a while after consecutive errors, so traffic goes to the healthy ones. Circuit breaking (connectionPool limits) caps how much a client sends to a service at all; above the limit the proxy fails requests immediately with flag UO instead of queueing them.

Why does a fault delay not trigger the timeout on the same route?

Envoy applies fault injection before routing, so a delay injected on a route is not counted by that route's timeout. To test timeouts, inject the fault at the dependency and set the timeout on the caller above it, which is exactly what happens in a real slowdown.

Where should timeouts be set in a call chain?

Shorter the deeper you go: if the web tier waits 9 seconds, the API it calls should give up after less, and the API's dependency sooner still. Otherwise outer layers give up while inner ones keep working on requests nobody is waiting for.

In an interview Mid

What is a retry storm, and how do you prevent one in a service mesh?

When every layer of a call chain retries, one failing service gets the product of all those retries: with Istio's default of 2 attempts at three layers, a request can become up to 27 calls to the bottom service, which makes it fail harder. To prevent it I retry at one layer, usually the one closest to the failing dependency, and set retries.attempts: 0 in the VirtualServices of the other layers. Timeouts get shorter the deeper the call goes, so outer layers do not give up while inner ones still work. Outlier detection and connection pool limits stop sending traffic to endpoints that are already failing.

Also asked: What is the difference between a VirtualService and a DestinationRule? · How would you test a timeout with fault injection? · What does outlier detection do?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.