The problem it solves
Fifty services, written by ten teams in four languages. Security wants every call encrypted and every service to prove who it is. SRE wants the same retries and timeouts everywhere, and a log line for every request between services. Asking each team to build that into each app never ends.
A service mesh moves it out of the apps into the network: a proxy next to every workload intercepts its traffic and does it for them.
What you need to know already: pods and their containers, init containers and native sidecars (15.12), Services and ClusterIPs (16.1-16.3), TLS and mutual TLS (9.15-9.19), the edge proxies of this chapter, iptables (16.3).
Data plane and control plane
control plane (istiod)
config (xDS) + certificates (CA)
┌──────────────┼──────────────┐
pod web v v v pod cart
[ app ]──localhost──[ Envoy ]═══mTLS═══[ Envoy ]──localhost──[ app ]
data plane: one proxy per pod (the "sidecar")
- The data plane is the proxies (Envoy, in Istio) that carry the traffic.
- The control plane (istiod) watches Kubernetes (Services, endpoints, the mesh's own CRDs), computes each proxy's configuration and pushes it over xDS (Envoy's config API: CDS clusters, LDS listeners, RDS routes, EDS endpoints). It is also the CA that issues every workload a short-lived certificate.
The app does not change: it still calls http://cart. iptables rules in the pod's network namespace redirect its outgoing traffic to the local Envoy (port 15001) and incoming traffic to Envoy first (port 15006).
What you get
security mutual TLS on every hop, a cryptographic identity per workload
(spiffe://cluster.local/ns/shop/sa/web), policy on who may call what
reliability retries, timeouts, circuit breaking, outlier detection - per route,
without code; fault injection to test them
traffic weighted splits, header routing, mirroring - for east-west calls too
telemetry the same request metrics and access log line for every service,
whatever language it is written in
What it costs
Be honest about this; interviewers check whether you know it:
- Resources: an Envoy per pod - typically 40-100 MiB of memory and some CPU each. 300 pods = 300 proxies.
- Latency: two extra proxy hops per call (client sidecar, server sidecar), usually around a millisecond each - more under load or with big configs.
- Complexity: a second networking layer to debug. "Is it the app, the sidecar, the policy, or the control plane?" is a new question at 3 am.
- Upgrades: the proxies are in every pod; a mesh upgrade means restarting every workload (rolling), and the control plane and proxies must stay within supported version skew.
- Gotchas: startup ordering (the app starting before its proxy), Jobs that never finish because the sidecar keeps running (solved by native sidecars), protocol detection on badly named ports.
Shapes of mesh
sidecar a proxy container in every pod Istio (classic), Linkerd
ambient a node-level L4 proxy (ztunnel) + optional Istio ambient (GA since 1.24)
per-namespace L7 proxies (waypoints)
proxyless the app's gRPC library talks xDS itself gRPC + istiod / Traffic Director
eBPF/node the CNI does identity and encryption Cilium service mesh
This chapter uses Istio in sidecar mode: it is the most deployed shape and the one whose behaviour you can see in every pod. The last lesson compares the others.
When not to
A handful of services, one language with a good client library, no compliance requirement for mTLS everywhere, a small platform team: a mesh is likely more risk than value. NetworkPolicies (16.29), a Gateway at the edge, and timeouts in the app's HTTP client get you most of the way. Adopt a mesh for a concrete need (mTLS for an audit, uniform telemetry across many teams, safe canaries between services), not because it is in the architecture diagram of a conference talk.
In an interview: "A mesh puts a proxy next to every pod; the control plane configures them and issues identities. You get mTLS, retries/timeouts and uniform telemetry without changing apps, and you pay in memory per pod, latency per hop, and operational complexity."
What you can now do:
- explain data plane vs control plane, and what xDS and the mesh CA do
- list what a mesh gives and what it costs, with numbers
- argue for or against a mesh for a given team