OnCallReady

Lesson 34.38 · Kubernetes: Ingress, Gateway API & Service Mesh · 9 min read

Ambient mode, Linkerd, and when not to use a mesh

In plain words

Imagine replacing the personal assistant at every desk with two kinds of shared staff: one security guard per floor who checks badges and seals envelopes for everybody, and, only for the departments that need it, a specialist who reads the letters and sorts them by content. Fewer people, and you pay for the specialist only where you use one.

ambient mode is that design for Istio: a ztunnel per node does L4 work (mTLS and identity) for every pod on it, and an optional waypoint proxy per namespace or Service does L7 work (routing, retries, HTTP policies). The lesson also compares Linkerd (smaller, simpler) and Cilium (mesh features in the network layer), introduces GAMMA (Gateway API routes for mesh traffic) and ends with how to decide whether to run a mesh at all.

Ambient: a mesh without sidecars

Sidecars cost memory per pod and a restart for every mesh upgrade. Istio's ambient mode (GA since 1.24, November 2024) splits the proxy in two:

ztunnel     one per NODE (a DaemonSet, written in Rust): L4 only - mTLS, identity,
            L4 authorization, TCP metrics. Every pod of an ambient namespace gets it.
waypoint    one per namespace or service, optional, a normal Envoy Deployment:
            L7 - HTTP routing, retries, L7 authorization - only where you need it.
kubectl label namespace shop istio.io/dataplane-mode=ambient    # no restart, no sidecars
istioctl waypoint apply -n shop --enroll-namespace              # add L7 when needed

What changes for you: no injection, no 2/2 pods, no restart to join or upgrade; the CNI redirects traffic to ztunnel. What stays: the same CRDs (AuthorizationPolicy, PeerAuthentication), and for L7 the same Envoy - in the waypoint. Debugging moves from kubectl logs pod -c istio-proxy to ztunnel and waypoint logs. Sidecars remain fully supported; both modes can share one mesh.

What you need to know already: the Istio lessons of this chapter, DaemonSets (15.20), Gateway API (this chapter).

Linkerd

The other long-lived mesh, and the simplest:

                Istio (sidecar)                     Linkerd
proxy           Envoy (C++), very configurable      linkerd2-proxy (Rust), purpose-built
footprint       ~50-100 MiB per proxy               ~10-20 MiB per proxy
mTLS            on (PERMISSIVE by default)          on by default, nothing to configure
config API      VirtualService, DestinationRule...  mostly Gateway API routes + its own policies
features        everything (faults, mirrors,        fewer knobs: retries, timeouts, splits,
                Wasm, ext authz, multicluster)      golden metrics, very good defaults
releases        CNCF graduated, open source         open-source edge releases only since 2024;
                stable releases                     stable builds come from vendors (Buoyant)

Linkerd's pitch: fewer features, fewer ways to break it. Istio's: whatever you need, it can do - at the cost of a bigger surface.

Cilium offers mesh features in the CNI itself (eBPF for L4, an Envoy per node for L7), attractive if Cilium is already your CNI.

One API for the mesh too: GAMMA

Gateway API's GAMMA initiative lets an HTTPRoute attach to a Service instead of a Gateway (parentRefs: [{kind: Service, name: reviews}]): mesh routing rules in the same API as the edge. Istio and Linkerd both support it, so a team can learn one routing API for north-south and east-west.

Deciding

Questions to ask before adopting a mesh:

need                                   without a mesh
mTLS everywhere for an audit           cert-manager + app TLS (hard at scale) -> a mesh wins
L4 segmentation                        NetworkPolicy is enough
uniform golden-signal metrics          client libraries / OpenTelemetry in each app
retries + timeouts                     the HTTP client library (often better: it knows idempotency)
canaries between services              Argo Rollouts / Gateway API at the edge

And who will run it: a mesh is a second network, with its own upgrades, its own CVEs, its own 3 am failures (istiod down = nothing can roll out; a bad AuthorizationPolicy = a self-inflicted outage). If nobody owns it, do not install it. If you do: start PERMISSIVE, one namespace, access logs on, istioctl analyze in CI, and grow from there.

What you can now do:

Why it helps

"Sidecar or ambient?" and "Istio or Linkerd?" are standard interview questions now, and the honest answer depends on trade-offs you can only explain if you know how each works: what a node-level proxy changes for resource use, upgrades and blast radius, and which features need an L7 proxy.

The decision lesson matters even more in practice. Many teams adopt a mesh because it sounds modern and then struggle to run it. Being able to say "these are our problems, this is what a mesh would cost, and this smaller tool solves what we need" is a senior habit worth building early.

Commands in this lesson

istioctl

FAQ

Does ambient mode remove all proxies from my pods?

Yes, pods run without a sidecar. Their traffic is redirected to the ztunnel on the node, which handles mTLS and L4 policy. L7 features need a waypoint, which is a normal Envoy Deployment that traffic for a namespace or Service passes through. Pods opt in with a namespace label instead of being re-created.

Can sidecars and ambient run in one mesh?

Yes. Istio supports both modes together, so a team can move namespaces one at a time. The same CRDs (PeerAuthentication, AuthorizationPolicy, routing) apply, though L7 rules in ambient need a waypoint to enforce them. That makes a gradual move practical.

Why choose Linkerd over Istio?

Linkerd focuses on a smaller feature set (mTLS, retries, timeouts, metrics, traffic splitting) with its own lightweight proxy, and is usually simpler to run and upgrade. Istio offers more features and extension points. If you need only the core, the simpler tool is often the better choice.

What is GAMMA?

The Gateway API initiative for mesh traffic: an HTTPRoute whose parent is a Service instead of a Gateway describes routing for calls to that Service inside the mesh. It means the same route API for the edge and for east-west traffic, and implementations such as Istio and Linkerd support it.

What is the blast radius of a ztunnel failure?

One ztunnel serves every ambient pod on its node, so a problem with it affects that whole node's mesh traffic, where a broken sidecar affects one pod. In exchange, there are far fewer proxies to run, upgrade and pay for. That trade-off is the heart of the sidecar versus ambient question.

In an interview Mid

What is Istio ambient mode, and how does it differ from sidecars?

In ambient mode there is no proxy in the pod. A ztunnel on each node handles L4 for all ambient pods on it: mTLS, workload identity and L4 authorization. L7 features (HTTP routing, retries, header-based policies) come from an optional waypoint proxy per namespace or Service. Compared with sidecars that means much less memory, no pod restarts to join or upgrade, and paying for L7 only where it is used; the costs are a node-wide blast radius for ztunnel and an extra hop through the waypoint for L7 traffic. Both modes can run in one mesh, so migration can be gradual.

Also asked: How does Linkerd differ from Istio? · What is GAMMA in Gateway API? · How would you decide whether a team needs a service mesh at all?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.