After the v2 rollout of payments, every call from web to it fails. The pods are Running and Ready, and the app logs show nothing, because no request reached them:
$ kubectl exec -n uh-lab deploy/web -- curl -s -w " %{http_code}\n" http://payments/api/charge
no healthy upstream 503no healthy upstream is not an error your app writes. It is Envoy, the sidecar proxy next to web, saying it had nowhere to send the request. Its access log says so in two letters:
$ kubectl logs -n uh-lab deploy/web -c istio-proxy --tail=1
[2026-09-22T20:00:04.500Z] "GET /api/charge HTTP/1.1" 503 UH no_healthy_upstream - "-" 0 19 0 - "-" "curl/8.16.0" "d0d33277-9f54-ae74-9f94-965b3d0f5bf0" "payments" "-" outbound|80|v2|payments.uh-lab.svc.cluster.local - 10.96.204.113:80 10.244.1.53:54145 - -What is happening: two proxies, one log line each
In sidecar mode, every request in the mesh passes through two Envoys: the caller's (outbound) and the server's (inbound). With access logging on (meshConfig.accessLogFile: /dev/stdout, or a Telemetry resource per namespace), each one writes one line per request. The fields right after the status code are the ones that matter:
503 the status code sent back to the caller
UH RESPONSE_FLAGS: why, in Envoy's words ("-" = nothing special)
no_healthy_upstream RESPONSE_CODE_DETAILS: who made the code
...
outbound|80|v2|payments.uh-lab... UPSTREAM_CLUSTER: direction|port|subset|serviceThe rule that saves hours: a 503 from the app has flag - and details via_upstream. A 503 that the proxy made itself has a flag. Look at the flag before you look at the app.
The flags you will meet most often:
| flag | meaning | usual cause |
|---|---|---|
UH | no healthy upstream | no endpoints: a subset matching no pod, all endpoints ejected, a Service with no ready pods |
UF | upstream connection failure | nothing listens on the target port, or a TLS/mTLS mismatch |
URX | retries exhausted | the upstream kept failing (5xx or connect errors) past the retry limit |
NR | no route | no VirtualService rule matches this request |
NC | no cluster | the route points at a subset no DestinationRule defines |
UT | upstream request timeout | the route's timeout expired: a 504 |
UO | upstream overflow | a circuit breaker (connectionPool limits) said no |
UC / DC | connection closed by upstream / by the caller | resets, the caller's own timeout |
FI / DI | fault injected (abort / delay) | a VirtualService fault you forgot about |
Recent Envoy versions can also write them out in full (%RESPONSE_FLAGS_LONG%: NoHealthyUpstream, UpstreamRetryLimitExceeded...) if you customise the log format.
Diagnosis, flag by flag
UH: what endpoints does the proxy have?
istioctl proxy-config (pc) shows what one proxy was told. The cluster from the log line exists:
$ istioctl pc clusters deploy/web -n uh-lab --fqdn payments
SERVICE FQDN PORT SUBSET DIRECTION TYPE DESTINATION RULE
payments.uh-lab.svc.cluster.local 80 - outbound EDS payments.uh-lab
payments.uh-lab.svc.cluster.local 80 v2 outbound EDS payments.uh-labBut it has no endpoints:
$ istioctl pc endpoints deploy/web -n uh-lab --cluster "outbound|80|v2|payments.uh-lab.svc.cluster.local"
ENDPOINT STATUS OUTLIER CHECK CLUSTERThe DestinationRule's v2 subset selects version: v2. The pods say:
$ kubectl get pods -n uh-lab --show-labels | grep payments
payments-v2-hhtm2hs85v-729qg 2/2 Running 0 10m app=payments,version=v2.0,pod-template-hash=hhtm2hs85v,security.istio.io/tlsMode=istio,service.istio.io/canonical-name=payments,service.istio.io/canonical-revision=v2.0
payments-v2-hhtm2hs85v-m5hbb 2/2 Running 0 10m app=payments,version=v2.0,pod-template-hash=hhtm2hs85v,security.istio.io/tlsMode=istio,service.istio.io/canonical-name=payments,service.istio.io/canonical-revision=v2.0v2 is not v2.0. Fix the subset's label, and the same curl returns 200. An empty endpoint list means a selector problem (here, or in the Service itself: a Service with no endpoints). Endpoints listed with OUTLIER CHECK FAILED mean outlier detection threw them out.
UF and URX: connect failures and retries
Here a backend outside the mesh (1/1 READY, no sidecar) sits behind a Service whose targetPort is 9999, where nothing listens:
$ kubectl exec -n d-fl-wallet deploy/client -- curl -s -w ' %{http_code}\n' http://wallet/data
upstream connect error or disconnect/reset before headers. retried and the latest reset reason: connection failure 503
$ kubectl logs -n d-fl-wallet deploy/client -c istio-proxy --tail=1
[2026-09-22T20:00:31.900Z] "GET /data HTTP/1.1" 503 UF,URX upstream_reset_before_response_started{connection_failure} - "delayed_connect_error:_Connection_refused" 0 114 53 - "-" "curl/8.16.0" "3ade11b0-fceb-be01-abcf-1e04622a74f3" "wallet" "10.244.1.90:9999" outbound|80||wallet.d-fl-wallet.svc.cluster.local 10.244.1.53:41335 10.96.235.96:80 10.244.1.53:34064 - defaultUF: the TCP connect to 10.244.1.90:9999 was refused (the transport failure reason says so). URX: Istio's default retry policy (2 retries, on connect failures and 503s, among others) tried again and gave up. The 53 ms duration is three attempts with a short backoff. UPSTREAM_HOST names the pod and port it tried, and the port is the first thing to check.
When both ends have sidecars, the client usually logs URX with via_upstream instead, because it connected fine to the server's sidecar. The UF is then in the server's proxy log, which failed to reach the app on localhost. Always read both lines, matched by the x-request-id.
URX without UF means the app itself kept answering 5xx:
$ kubectl logs -n d-fl-wallet deploy/client -c istio-proxy --tail=1
[2026-09-22T20:00:06.000Z] "GET /data HTTP/1.1" 503 URX via_upstream - "-" 0 74 57 2 "-" "curl/8.16.0" "d0d33277-9f54-ae74-9f94-965b3d0f5bf0" "wallet" "10.244.2.25:8080" outbound|80||wallet.d-fl-wallet.svc.cluster.local 10.244.1.53:57173 10.96.235.96:80 10.244.1.53:49902 - defaultvia_upstream: the app sent the 503 ("error": "wallet is overloaded" in the body), three times. The fix is in the app, or in its capacity. Retries on an overloaded service multiply its load, which is how retry storms start.
NR, NC, UT: the routing config said so
Three client-side lines, from three namespaces (kubectl logs deploy/client -c istio-proxy --tail=1 in each):
[2026-09-22T20:00:06.000Z] "GET /api/v2/list HTTP/1.1" 404 NR route_not_found - "-" 0 0 0 - "-" "curl/8.16.0" "d0d33277-9f54-ae74-9f94-965b3d0f5bf0" "rates" "-" - - 10.96.235.96:80 10.244.1.53:54145 - -
[2026-09-22T20:00:06.000Z] "GET /data HTTP/1.1" 503 NC cluster_not_found - "-" 0 0 0 - "-" "curl/8.16.0" "d0d33277-9f54-ae74-9f94-965b3d0f5bf0" "ledger" "-" - - 10.96.235.96:80 10.244.1.53:54145 - -
[2026-09-22T20:00:06.000Z] "GET /api/items HTTP/1.1" 504 UT response_timeout - "-" 0 24 1000 - "-" "curl/8.16.0" "d0d33277-9f54-ae74-9f94-965b3d0f5bf0" "fraud" "10.244.2.25:8080" outbound|80||fraud.d-fl-fraud.svc.cluster.local 10.244.1.53:41335 10.96.235.96:80 10.244.1.53:34064 - -No upstream host and no cluster in the first two: the request never left the proxy. In each case, kubectl get vs,dr -n <ns> shows the culprit. The NR VirtualService only has a rule for prefix: /admin. The NC one routes to subset: canary, which no DestinationRule defines. The UT one has timeout: 1s, and the duration field says exactly 1000 ms. A route timeout shorter than the service's real latency turns slow into broken. p99 latency tells you what the timeout should be.
Before reading YAML by hand, istioctl analyze -n <ns> catches many of these (a subset that is not defined, a host that does not exist).
Keeping it from coming back
- Turn access logs on where you run production traffic, at least for errors (a Telemetry
filteronresponse.code >= 400). Without them, every 503 starts as a guess. - Alert on flags.
istio_requests_totalcarries aresponse_flagslabel: a rate ofUHorNRabove zero is a config bug, never load. - Version labels from the deploy tool (the image tag), never typed by hand, so subsets and pods cannot drift apart.
- Know your default retries before an incident does: set
retriesexplicitly per route, and lower them for services that can overload. - The same "who said no" question applies at the edge. A 502 from nginx has its own error log to read, and a Gateway that ignores a route says why in its status conditions.
Practise it
In the chapter Kubernetes: Ingress, Gateway API & Service Mesh, the mission Read a 503 UH is the subset mismatch above, solved with proxy-config. The drill Read the response flag gives you a fresh broken namespace each round (UH, NC, NR, FI, UT, URX or an RBAC denial) to diagnose from the access log alone. The lesson Reading Envoy: the access log, response flags and proxy-config goes through every field of the log line.