OnCallReady

Lesson 34.37 · Kubernetes: Ingress, Gateway API & Service Mesh · 8 min read

Mesh telemetry: golden signals from every proxy

In plain words

Imagine every till in a supermarket chain reported the same three numbers every minute: how many customers paid, how many walked away annoyed, and how long the queue took. Head office would see which shop is struggling without asking any shop manager to install anything.

The mesh does that for services. Every proxy counts requests, errors and durations with the same labels (source, destination, response code), so you get the golden signals for every service without changing the app. In this lesson you read those metrics, measure latency yourself with fortio and its target 50% / 90% / 99% lines (in seconds), turn access logs on per namespace with the Telemetry API, and see why tracing still needs the app's help.

Telemetry you did not write

Every proxy measures every request: count, status code, latency, bytes, with labels for the source and destination workload. Because the same proxy sits next to every service, you get the same metrics for the Go service, the Java service and the vendor black box.

What you need to know already: the Envoy lesson before this one, latency percentiles (0.4), golden signals (0.5).

The standard metrics

istio_requests_total{reporter, source_workload, source_workload_namespace,
                     destination_workload, destination_service, response_code,
                     response_flags, connection_security_policy, ...}
istio_request_duration_milliseconds_bucket{...}     a histogram: percentiles
istio_request_bytes / istio_response_bytes
istio_tcp_connections_opened_total, istio_tcp_sent_bytes_total   (TCP traffic)

reporter="source" = counted by the client sidecar, "destination" = by the server sidecar. The two differ exactly when the network or a policy lies between them: a request the client counted with flag UF never reached the server.

Each pod serves them on port 15020 (/stats/prometheus, merged with the app's own metrics) - that is what the injected prometheus.io/scrape annotations point at.

Later (Ch 27): Prometheus scrapes port 15020 of every pod, and PromQL turns istio_requests_total into per-service error rates and istio_request_duration_milliseconds_bucket into p99 latency.

Later (Ch 29): dashboards and service graphs (Grafana, Kiali) are built from exactly these labels, and distributed traces need the apps to forward the trace headers.

Measuring it yourself: fortio

fortio is the load generator Istio's own docs use. Run it from a pod in the mesh, so its calls go through a sidecar:

$ kubectl exec deploy/fortio -c fortio -- fortio load -c 2 -qps 0 -n 40 -loglevel Warning http://ratings:80/api
...
# target 50% 0.0034
# target 90% 0.0081
# target 99% 2.0021
...
Code 200 : 40 (100.0 %)

-c connections (concurrency), -qps 0 as fast as possible, -n calls. Durations are in seconds: p50 3.4 ms, p99 2 s. A p99 two orders of magnitude above the median is the signature of a slow minority - 1 request in 50 hitting a 2 s fault, or a GC pause, or one bad pod.

Access logs per namespace: the Telemetry API

meshConfig.accessLogFile is all-or-nothing. The Telemetry resource scopes it (and metrics and tracing) by namespace or workload:

apiVersion: telemetry.istio.io/v1
kind: Telemetry
metadata: {name: access-logs, namespace: payments}
spec:
  accessLogging:
  - providers: [{name: envoy}]

In istio-system it applies to the whole mesh; a workload selector narrows it. Access logs cost disk and money at volume - many teams log only errors with a filter: {expression: "response.code >= 400"}.

Tracing needs the app

Envoy creates spans for every hop, but it cannot connect "the request web received" with "the request web then made to api" - only the app knows. Apps must copy the trace headers (traceparent, x-request-id, the x-b3-* set) from incoming to outgoing requests. Without that you get one-hop traces.

What you can now do:

Why it helps

Uniform metrics are one of the strongest reasons to run a mesh: every service, in every language, reports the same request rate, error rate and latency, so dashboards and alerts can be written once. Knowing where those numbers come from (the proxies, labelled by reporter, source and destination) is what lets you trust them, or spot when they disagree with the app.

Measuring percentiles yourself matters too: the median barely moves when a slow dependency hits half the requests, while p99 jumps to the delay. Reading fortio's output, and knowing why tracing headers must be copied by the app, keeps you from drawing the wrong conclusion in an incident.

Commands in this lesson

kubectl

FAQ

What is the reporter label?

Each request is counted by both proxies: reporter="source" by the client's sidecar and reporter="destination" by the server's. Pick one when you add numbers up, otherwise every request is counted twice; the destination view is the usual choice for a service's own error rate.

Why look at p99 instead of the average?

Averages hide the slow tail. If one request in a hundred takes two seconds, the average barely moves, but a page that makes a hundred calls is almost always slow. p99 is what your unluckiest users feel, and it moves first when a dependency starts struggling.

Why does tracing need changes in the app if the proxies see everything?

Each proxy sees one hop. To join the hops into one trace, the trace headers from the incoming request must be copied onto the outgoing calls the app makes, and only the app knows which outgoing call belongs to which incoming request. Without that, each hop starts a new trace.

What does the Telemetry API control?

Access logging, metrics and tracing settings per mesh, namespace or workload, without editing the global mesh config. A typical use is turning access logs on for one noisy namespace while debugging, then removing the Telemetry object afterwards.

Is fortio a mesh tool?

It is a load-testing tool from the Istio project, often used with meshes because it prints latency percentiles and response codes. It is not part of the mesh; you run it as a pod, from inside the mesh, to send controlled load and see what users would experience.

In an interview Mid

What telemetry does a service mesh give you without changing the applications?

Every proxy reports the same request metrics for every service: a request count with labels for source, destination and response code, and request duration, so I get rate, errors and latency per service and per caller in any language. Access logs per request come from the proxies too, and the Telemetry API can turn them on per namespace. What the mesh cannot do alone is tracing across hops: the app must copy the trace headers onto its outgoing calls. To check latency myself I run fortio and read the target 99% line, in seconds, because the tail moves before the median.

Also asked: What labels do the standard Istio request metrics carry? · Why do p99 and the median tell different stories? · Why must applications propagate trace headers?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.