OnCallReady

Lesson 34.27 · Kubernetes: Ingress, Gateway API & Service Mesh · 9 min read

What a service mesh gives you, and what it costs

In plain words

Imagine every employee in a company had to handle their own post security: check every letter's sender, seal every envelope, keep a log of what they sent, and resend what got lost. Nobody would get any real work done. So the company puts a post clerk next to every desk who does all of that, and a central office tells all the clerks the current rules.

A service mesh is that arrangement for Services. A proxy next to each pod (the data plane) handles encryption, identity, retries, timeouts and measurement for every call; a control plane pushes configuration to all the proxies over xDS and acts as the CA that gives each workload its certificate. The apps do not change. The price is a proxy per pod: memory, a little latency, and one more system to run.

The problem it solves

Fifty services, written by ten teams in four languages. Security wants every call encrypted and every service to prove who it is. SRE wants the same retries and timeouts everywhere, and a log line for every request between services. Asking each team to build that into each app never ends.

A service mesh moves it out of the apps into the network: a proxy next to every workload intercepts its traffic and does it for them.

What you need to know already: pods and their containers, init containers and native sidecars (15.12), Services and ClusterIPs (16.1-16.3), TLS and mutual TLS (9.15-9.19), the edge proxies of this chapter, iptables (16.3).

Data plane and control plane

                         control plane (istiod)
                   config (xDS) + certificates (CA)
                 ┌──────────────┼──────────────┐
pod web          v              v              v        pod cart
[ app ]──localhost──[ Envoy ]═══mTLS═══[ Envoy ]──localhost──[ app ]
             data plane: one proxy per pod (the "sidecar")

The app does not change: it still calls http://cart. iptables rules in the pod's network namespace redirect its outgoing traffic to the local Envoy (port 15001) and incoming traffic to Envoy first (port 15006).

What you get

security       mutual TLS on every hop, a cryptographic identity per workload
               (spiffe://cluster.local/ns/shop/sa/web), policy on who may call what
reliability    retries, timeouts, circuit breaking, outlier detection - per route,
               without code; fault injection to test them
traffic        weighted splits, header routing, mirroring - for east-west calls too
telemetry      the same request metrics and access log line for every service,
               whatever language it is written in

What it costs

Be honest about this; interviewers check whether you know it:

Shapes of mesh

sidecar        a proxy container in every pod                 Istio (classic), Linkerd
ambient        a node-level L4 proxy (ztunnel) + optional     Istio ambient (GA since 1.24)
               per-namespace L7 proxies (waypoints)
proxyless      the app's gRPC library talks xDS itself        gRPC + istiod / Traffic Director
eBPF/node      the CNI does identity and encryption           Cilium service mesh

This chapter uses Istio in sidecar mode: it is the most deployed shape and the one whose behaviour you can see in every pod. The last lesson compares the others.

When not to

A handful of services, one language with a good client library, no compliance requirement for mTLS everywhere, a small platform team: a mesh is likely more risk than value. NetworkPolicies (16.29), a Gateway at the edge, and timeouts in the app's HTTP client get you most of the way. Adopt a mesh for a concrete need (mTLS for an audit, uniform telemetry across many teams, safe canaries between services), not because it is in the architecture diagram of a conference talk.

In an interview: "A mesh puts a proxy next to every pod; the control plane configures them and issues identities. You get mTLS, retries/timeouts and uniform telemetry without changing apps, and you pay in memory per pod, latency per hop, and operational complexity."

What you can now do:

Why it helps

Mesh questions are a staple of platform and SRE interviews, and the strongest answers are honest about the costs, not only the features. Knowing what a mesh gives (mTLS without app changes, uniform retries and timeouts, the same metrics for every service) and what it costs (resources, latency, complexity, upgrades) lets you argue for or against one in a design review.

It also explains the rest of the chapter: every Istio object you will write is configuration for those proxies, and every failure you will debug is a proxy doing exactly what it was told. If you understand the split between data plane and control plane, the commands make sense.

Commands in this lesson

istioctl

FAQ

Does a mesh replace Kubernetes Services?

No. The proxies still use Services and their endpoints to know where workloads are. The mesh adds behaviour on top: which version gets traffic, how retries work, who may call whom, and encryption. Remove the mesh and Service-to-Service traffic still flows, just without those features.

How much does a sidecar cost?

Each proxy uses memory (tens of megabytes, more in large meshes because every proxy knows about many services) and some CPU, and every call goes through two proxies, adding latency in the order of a millisecond. Multiplied by hundreds of pods, that is real money and a real tail-latency budget.

What is the control plane in Istio?

istiod. It watches Kubernetes Services, endpoints and Istio objects, turns them into Envoy configuration and sends it to every proxy over xDS, and it is the certificate authority that signs each workload's identity certificate. If istiod is down, proxies keep their last configuration, but new pods cannot get injected or receive config.

What does proxyless mean?

Some gRPC clients can talk xDS themselves, so they get routing and service discovery from the control plane without a sidecar. It removes the proxy's cost but puts the logic back into the application's libraries, and it does not cover every language or feature.

When is a mesh a bad idea?

When you have a handful of services, no requirement for mTLS between them, and no one with time to run it. A mesh adds a moving part to every request; if the problems it solves are not your problems, an ingress, NetworkPolicies and good client libraries may be enough.

In an interview Mid

What does a service mesh give you, and what does it cost?

A mesh puts a proxy next to every workload, the data plane, and a control plane configures those proxies and issues each workload an identity certificate. Without changing the apps I get mTLS and identity-based authorization between services, consistent retries, timeouts and traffic splitting, and the same request metrics for every service. The costs are real: memory and CPU per pod, extra latency per hop, a complex system with its own upgrades and failure modes, and debugging that now includes the proxies. I would recommend one when there are many services and a security or reliability requirement it solves, not by default.

Also asked: What is the difference between the data plane and the control plane? · What is xDS, and what does the control plane send over it? · When would you decide against a service mesh?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.