OnCallReady

Lesson 34.36 · Kubernetes: Ingress, Gateway API & Service Mesh · 11 min read

Reading Envoy: the access log, response flags and proxy-config

In plain words

Imagine every delivery van in a city keeps a logbook: time, address, what was delivered, how long it took, and a short code when something went wrong ("no one home", "road closed", "gave up after two tries"). If a parcel goes missing, you do not guess; you read the logbooks of the two vans that handled it.

Every Envoy proxy in the mesh writes one access log line per request, and a request between two pods is logged twice: once by the client's proxy (outbound) and once by the server's (inbound). The response flag is the short code: UH no healthy upstream, NR no route, UO overflow, UT timeout, URX retries exhausted, DC the client went away. istioctl proxy-config shows what one proxy was told: clusters, listeners, routes, endpoints.

One line per request, per proxy

With access logging on, every proxy writes one line per request it handled. A call from curl to httpbin in the mesh is logged twice: by the client sidecar (outbound) and by the server sidecar (inbound).

What you need to know already: the previous lessons of this part, the ingress-nginx access log (16.19) for comparison, HTTP status codes (9.22).

[2026-10-07T10:12:01.400Z] "GET /status/503 HTTP/1.1" 503 URX via_upstream - "-" 0 0 54 1 "-" "curl/8.16.0" "963585ef-092e-cf6b-b305-fb8ddaf44fa0" "httpbin:8000" "10.244.2.229:8080" outbound|8000||httpbin.demo.svc.cluster.local 10.244.2.192:45011 10.96.204.113:8000 10.244.2.192:37740 - default

Istio's default format, field by field:

[2026-10-07T10:12:01.400Z]            START_TIME
"GET /status/503 HTTP/1.1"            method, path, protocol
503                                   RESPONSE_CODE sent to the downstream (the caller)
URX                                   RESPONSE_FLAGS: why, in Envoy's words ("-" = nothing special)
via_upstream                          RESPONSE_CODE_DETAILS: who made the code (the upstream answered)
-                                     CONNECTION_TERMINATION_DETAILS
"-"                                   UPSTREAM_TRANSPORT_FAILURE_REASON (TLS errors, connect errors)
0 0                                   BYTES_RECEIVED, BYTES_SENT
54                                    DURATION (ms, the whole request incl. retries)
1                                     X-ENVOY-UPSTREAM-SERVICE-TIME (ms the last upstream try took)
"-"  "curl/8.16.0"                    X-Forwarded-For, User-Agent
"963585ef-..."                        X-REQUEST-ID: the same id on every hop - grep for it
"httpbin:8000"                        authority (Host)
"10.244.2.229:8080"                   UPSTREAM_HOST: the pod that was tried last
outbound|8000||httpbin.demo.svc...    UPSTREAM_CLUSTER: direction|port|subset|host
10.244.2.192:45011                    UPSTREAM_LOCAL_ADDRESS
10.96.204.113:8000                    DOWNSTREAM_LOCAL_ADDRESS (what the app dialled: the ClusterIP)
10.244.2.192:37740                    DOWNSTREAM_REMOTE_ADDRESS (the caller)
-                                     REQUESTED_SERVER_NAME (SNI: outbound_.8000_._.httpbin... on inbound mTLS)
default                               ROUTE_NAME (the VirtualService route's name, or default)

Duration 54 ms for a 1 ms upstream answer: two retries with 25 ms backoff. URX says the retries ran out. Default retries, visible in one field.

Response flags

UH   no healthy upstream: no endpoints, all ejected, or a subset matching no pod   503 "no healthy upstream"
UF   upstream connection failure: connect refused / timed out                     503 "upstream connect error or disconnect/reset before headers. reset reason: connection failure"
UO   upstream overflow: a circuit breaker (connectionPool) said no                503
NR   no route: no VirtualService rule matched, or (inbound) no filter chain       404 / reset
URX  retries (or connect attempts) exhausted - usually with UF or after 5xx
UT   upstream request timeout: the route's timeout (or perTryTimeout) expired     504 "upstream request timeout"
UC   upstream connection termination: the server closed it (mTLS mismatch...)     503
DC   downstream connection termination: the CALLER went away (its timeout)
DI   a fault delay was injected
FI   a fault abort was injected                                                  "fault filter abort"
NC   no cluster: the route points at a host or subset istiod never sent           503
RL   rate limited locally                                                        429

The pairing that saves hours: a 503 from the app has flag "-" and via_upstream; a 503 the proxy made itself has a flag. Look at the flag before you look at the app.

What one proxy was told: proxy-config

Logs show what happened; istioctl proxy-config (pc) shows the configuration that made it happen, for one proxy:

istioctl pc clusters  deploy/web -n shop --fqdn reviews     SERVICE FQDN, PORT, SUBSET, DIRECTION, TYPE, DESTINATION RULE
istioctl pc endpoints deploy/web -n shop --cluster "outbound|9080|v2|reviews.shop.svc.cluster.local"
                                                             ENDPOINT, STATUS, OUTLIER CHECK, CLUSTER
istioctl pc routes    deploy/web -n shop --name 9080         NAME, VHOST NAME, DOMAINS, MATCH, VIRTUAL SERVICE
istioctl pc listeners deploy/web -n shop --port 9080         ADDRESSES, PORT, MATCH, DESTINATION
istioctl pc secret    deploy/web -n shop                     the workload certificate

The cluster name format outbound|9080|v2|reviews.shop.svc.cluster.local appears everywhere: logs, endpoints, stats. A subset in a VirtualService that no DestinationRule defines has no cluster in pc clusters; a subset whose labels match nothing has a cluster with no endpoints.

Envoy's own counters

The admin API listens on localhost:15000 inside the pod. The proxy image has no curl, so use pilot-agent:

$ kubectl exec deploy/web -c istio-proxy -- pilot-agent request GET stats | grep 'reviews.*upstream_rq_retry'
cluster.outbound|9080||reviews.shop.svc.cluster.local.upstream_rq_retry: 112

upstream_rq_retry, upstream_rq_pending_overflow (circuit breaker), upstream_rq_timeout, outlier_detection.ejections_active - each one a behaviour from the previous lesson, counted. Port 15020 serves the merged metrics of the app and the proxy at /stats/prometheus (istio_requests_total with source, destination, response code, flags and connection_security_policy).

What you can now do:

Why it helps

When a mesh misbehaves, the access log line and its response flag usually name the problem directly, and the proxy that wrote it tells you which hop. That is much faster than reading YAML and guessing which object is wrong.

proxy-config closes the loop: if the log says NC (no cluster) or the endpoints list is empty (no endpoints), you can see exactly which cluster or subset the proxy lacks, and from that which VirtualService or DestinationRule to fix. Envoy's own counters, through pilot-agent request GET stats, show retries, overflows and ejections that never appear in your app's logs.

Commands in this lesson

kubectl

FAQ

Why are there two log lines for one request?

The client's sidecar logs the outbound call and the server's sidecar logs the inbound one. Comparing them tells you where time went and where the failure started: a 503 UO only on the client side means the client's circuit breaker refused it before it ever left the pod.

How do I read the cluster name outbound|80|v2|reviews.shop.svc.cluster.local?

Direction, port, subset and the service's full name. An empty subset (outbound|80||...) means no subset. You will see these names in access logs, in istioctl proxy-config clusters and in stats, so learning the shape once pays off everywhere.

What do UF and UC mean?

UF is upstream connection failure: the proxy could not connect to the endpoint at all (often a crashed pod or a wrong port). UC is upstream connection termination: the connection broke after it was established. Both point at the backend or the network rather than at mesh configuration.

Are access logs on by default?

It depends on the installation. With meshConfig.accessLogFile=/dev/stdout every proxy logs to its container's output, so kubectl logs <pod> -c istio-proxy shows them. The Telemetry API can turn them on per namespace or workload instead of mesh-wide.

What does proxy-config endpoints add over kubectl get endpoints?

It shows the endpoints as one specific proxy sees them, with health and outlier status. An endpoint that Kubernetes lists as ready can still be marked FAILED by that proxy's outlier detection, which explains why one client avoids a pod that others use.

In an interview Mid

A call between two services fails with a 503 in the mesh. How do you find the cause?

I read the access log of the client's sidecar with kubectl logs <pod> -c istio-proxy and look at the response flag. UH means no healthy upstream: the subset or Service has no ready endpoints, which istioctl proxy-config endpoints confirms. NC means no cluster: the route points at a subset no DestinationRule defines. UO means a connection pool limit tripped, URX that retries were used up on a failing backend, UF that it could not connect at all. Then I check the server's inbound line: if there is none, the request never arrived. The flag tells me which object to fix.

Also asked: Why does one request produce two access log lines in a mesh? · What does istioctl proxy-config show you? · How would you see whether a proxy is retrying requests?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.