Two objects, two moments
Istio's traffic API has two halves, applied at two moments of a request:
VirtualService ROUTING: for requests to this host, which destination?
(matches, weights, retries, timeouts, faults, rewrites, mirrors)
DestinationRule POLICY after routing: how to talk to that destination
(subsets = named groups of pods, load balancing, TLS,
connection pools, outlier detection)
What you need to know already: the Istio and mTLS lessons, Deployments with labels (15.6), canary releases with weights (16.27), HTTP 5xx codes and timeouts (9.22-9.23), and the edge timeouts of this chapter.
apiVersion: networking.istio.io/v1
kind: DestinationRule
metadata: {name: reviews, namespace: shop}
spec:
host: reviews # the Service (short name = this namespace)
subsets:
- name: v1
labels: {version: v1} # pods of the Service with this label
- name: v2
labels: {version: v2}
---
apiVersion: networking.istio.io/v1
kind: VirtualService
metadata: {name: reviews, namespace: shop}
spec:
hosts: [reviews]
http:
- match:
- headers:
x-canary: {exact: "true"}
route:
- destination: {host: reviews, subset: v2}
- route:
- destination: {host: reviews, subset: v1}
weight: 90
- destination: {host: reviews, subset: v2}
weight: 10
The http list is ordered: the first match wins (unlike Gateway API's "most specific wins"). Put specific rules first, the catch-all last. A VS that matches nothing for a request gives a 404 with flag NR.
A route to a subset no DestinationRule defines is a 503 (flag NC, cluster not found) - istioctl analyze reports it as IST0101 before it hurts. A subset whose labels match no pod is a 503 no healthy upstream (flag UH).
Retries: there is a default
Even without any VirtualService, Istio retries: 2 attempts on connect-failure, refused-stream, unavailable, cancelled, retriable-status-codes - and the retriable status code is 503. So an upstream answering 503 is quietly called up to 3 times. Good for a flaky pod, bad when a dependency is overloaded.
http:
- route: [{destination: {host: inventory}}]
timeout: 3s # the whole call, all attempts included
retries:
attempts: 3
perTryTimeout: 1s # each attempt
retryOn: 5xx,gateway-error,connect-failure,reset
Rules of thumb:
timeoutmust be at leastattempts x perTryTimeout(plus backoff, 25 ms base) or the last attempts never happen.- Retry only idempotent requests (GET, PUT with the same body) - a retried POST can charge a card twice.
- Retry at one layer. web retries api 3 times, api retries inventory 3 times: one user request = up to 4 x 4 = 16 calls to inventory, exactly when it is already failing. That is a retry storm (the incident at the end of this part).
retries: {attempts: 0}switches them off for a route.
Fault injection
The client sidecar can break things on purpose, to test the retries, timeouts and fallbacks you just wrote:
http:
- fault:
delay: {percentage: {value: 50}, fixedDelay: 2s} # flag DI
abort: {percentage: {value: 10}, httpStatus: 503} # flag FI, body "fault filter abort"
route: [{destination: {host: ratings}}]
Faults happen in the client's sidecar before routing, so only callers in the mesh see them, and an injected delay does not count against that route's own timeout - the caller's timeout is what you are testing.
Outlier detection: take a bad pod out
One replica of five returns 500s (a bad node, a stuck cache). Load balancing keeps sending it 20% of traffic. Outlier detection ejects it:
spec:
host: inventory
trafficPolicy:
outlierDetection:
consecutive5xxErrors: 3 # 3 errors in a row from that host
interval: 10s # how often hosts are evaluated
baseEjectionTime: 30s # first ejection 30 s, then 60 s, 90 s...
maxEjectionPercent: 50 # never eject more than half (default 10%, at least one host)
Each client proxy decides on its own (there is no global view): in istioctl proxy-config endpoints the ejected host shows OUTLIER CHECK FAILED, and the stat outlier_detection.ejections_active counts it.
Circuit breaking: limit what you send
connectionPool caps how much one client proxy sends to a destination:
trafficPolicy:
connectionPool:
tcp: {maxConnections: 1}
http: {http1MaxPendingRequests: 1, maxRequestsPerConnection: 1}
Requests over the limit fail immediately with 503 and flag UO (upstream overflow) instead of piling up on a struggling service. It protects the server and frees the client's threads; the client must handle the fast 503 (a fallback, a cached answer). Load test it - fortio load -c 3 -qps 0 -n 30 URL from a pod in the mesh shows the overflow as Code 503.
In an interview: "VirtualService decides where a request goes - matches, weights, retries, timeouts, faults; DestinationRule decides how to talk to the destination - subsets, load balancing, connection pools, outlier detection. Istio retries 503s twice by default, which multiplies load across layers if every hop retries."
What you can now do:
- split traffic by weight and header with subsets
- set retries and a timeout that fit together, and know the default ones
- inject delays and aborts to test a caller
- eject a bad pod with outlier detection and cap load with a connection pool