OnCallReady

Lesson 34.28 · Kubernetes: Ingress, Gateway API & Service Mesh · 17 min read

Istio: install, injection and the control plane

In plain words

Imagine a new hire joins a company that gives every employee a personal assistant. The assistant is assigned on the first day, when the employee's desk is set up. People who were already there before the rule started do not get one until their desk is set up again.

Istio's sidecar works like that. You install the control plane with istioctl install and a profile, label a namespace istio-injection=enabled, and a webhook adds the Envoy proxy to every pod when they are created. Pods that already ran stay 1/1 until you restart them; new ones show 2/2. The proxy is a native sidecar (an init container that keeps running), and istioctl proxy-status and istioctl analyze tell you whether every proxy is in sync and whether your configuration makes sense.

Installing

The lab ships istioctl 1.31.1. It installs the control plane from a profile:

$ istioctl profile list
Istio configuration profiles:
    ambient
    default
    demo
    empty
    minimal
    remote

default = istiod + an ingress gateway; minimal = istiod only (use it with Gateway API for ingress); demo = more components and access logs on (labs, never production); ambient = the sidecar-less mode.

$ istioctl x precheck
✔ No issues found when checking the cluster. Istio is safe to install or upgrade!
$ istioctl install --set profile=default --set meshConfig.accessLogFile=/dev/stdout -y
✔ Istio core installed ⛵️
✔ Istiod installed 🧠
✔ Ingress gateways installed 🛬
✔ Installation complete

--set meshConfig.accessLogFile=/dev/stdout turns on Envoy's access log in every proxy - the single most useful debugging setting, off by default (the Telemetry API can do the same per namespace).

What landed:

istio-system/istiod                 the control plane: xDS server, CA, injection webhook
istio-system/istio-ingressgateway   an Envoy at the edge (default profile)
CRDs                                VirtualService, DestinationRule, PeerAuthentication,
                                    AuthorizationPolicy, Telemetry, ServiceEntry...
MutatingWebhookConfiguration        istio-sidecar-injector
ValidatingWebhookConfiguration      istio-validator-istio-system
GatewayClass istio                  Istio as a Gateway API implementation

Injection

Pods get a sidecar when they are created, from a mutating webhook, if their namespace asks for it:

kubectl label namespace shop istio-injection=enabled     # or istio.io/rev=<revision>
kubectl rollout restart deployment -n shop               # existing pods do not change by themselves

A pod can opt out with the label sidecar.istio.io/inject: "false". Pods that were running before the label stay 1/1 until they are recreated - the first thing istioctl analyze points out:

Warning [IST0103] (Pod shop/web-6d8f9c7b44-x2k8q) The pod shop/web-6d8f9c7b44-x2k8q is missing the Istio proxy. This can often be resolved by restarting or redeploying the workload.
Info [IST0102] (Namespace legacy) The namespace is not enabled for Istio injection. Run 'kubectl label namespace legacy istio-injection=enabled' to enable it, or 'kubectl label namespace legacy istio-injection=disabled' to explicitly mark it as not needing injection.

An injected pod shows 2/2 READY and two extra init containers:

initContainers:
- name: istio-init          # writes the iptables redirect rules, then exits
- name: istio-proxy         # restartPolicy: Always = a native sidecar (15.12)
  image: docker.io/istio/proxyv2:1.31.1

Since Istio 1.27-ish the proxy is a native sidecar whenever every node's kubelet is 1.33 or newer (ENABLE_NATIVE_SIDECARS=auto): it starts before the app containers and stops after them - which fixes the two classic sidecar bugs, "the app started before its proxy and failed its first calls" and "the Job never completes because the sidecar keeps running". On older clusters it was a normal container and you needed holdApplicationUntilProxyStarts.

The proxy's ports, worth knowing by heart:

15001  outbound capture (everything the app sends)
15006  inbound capture (everything sent to the pod)
15000  Envoy admin (localhost only: config_dump, stats, clusters)
15020  merged Prometheus metrics (app + Envoy), pilot-agent health
15021  health check (/healthz/ready)
15090  Envoy's own Prometheus stats

istio-init needs NET_ADMIN to write iptables rules. Clusters that forbid that use the Istio CNI plugin instead (the ambient profile installs it).

Is everything in sync?

$ istioctl proxy-status
NAME                                    CLUSTER      ISTIOD                    VERSION   SUBSCRIBED TYPES
curl-7c5d8f6b9d-wvm8l.demo              Kubernetes   istiod-5jb9vqzqx8-mpn9p   1.31.1    4 (CDS,LDS,EDS,RDS)
httpbin-bnhj652sp8-lt2bv.demo           Kubernetes   istiod-5jb9vqzqx8-mpn9p   1.31.1    4 (CDS,LDS,EDS,RDS)

One line per proxy: which istiod it is connected to, its version, the xDS types it subscribes to. -v 1 shows each type as SYNCED (3m2s), STALE (istiod sent config the proxy has not acknowledged - a bad push or an overloaded proxy) or NOT SENT. A pod missing from the list has no sidecar, or its proxy cannot reach istiod.

istioctl version              client, control plane and data plane versions (+ proxy count)
istioctl analyze -n shop      configuration problems; exit code 79 when it finds errors
istioctl x describe pod P     what applies to one pod: Service, DestinationRule, VirtualService, mTLS
istioctl proxy-config ...     what one Envoy was told (the lesson on reading Envoy)

Protocol selection

Envoy needs to know whether a port carries HTTP (it can then do L7 routing and metrics) or opaque TCP. Istio reads it from the Service port's appProtocol, else from the port name prefix (http-, http2-, grpc-, tcp-, tls-...), else it sniffs the first bytes. Name your ports; analyze reports IST0118 for names it cannot use.

The webhook fails closed

The injection webhook has failurePolicy: Fail. If istiod is down when a ReplicaSet creates a pod in an injected namespace, the pod is not created:

Warning  FailedCreate  replicaset-controller  Error creating: Internal error occurred: failed calling webhook "namespace.sidecar-injector.istio.io": failed to call webhook: Post "https://istiod.istio-system.svc:443/inject?timeout=10s": no endpoints available for service "istiod"

Running pods keep working (their proxies keep the last config), but nothing in those namespaces can scale or roll out. Run istiod with at least 2 replicas and a PodDisruptionBudget.

In an interview: "Injection happens at pod creation via a mutating webhook, triggered by the namespace label istio-injection=enabled, so existing pods need a rollout restart. I check with kubectl get pods (2/2), istioctl proxy-status, and istioctl analyze for IST0103."

What you can now do:

Why it helps

The most common "the mesh does nothing" ticket is a pod without a sidecar: the label was added after the pods started, or the pod opted out. Knowing that injection happens only at pod creation, and checking 2/2 and proxy-status, solves it in a minute.

The other failure this lesson prepares you for is quieter and worse: the injection webhook fails closed. If istiod is down, new pods in labelled namespaces cannot be created at all, so a rollout or a node failure turns into an outage. Recognising the FailedCreate event that names the webhook tells you to look at the control plane, not at your Deployment.

Commands in this lesson

istioctl

FAQ

Why are my pods still 1/1 after labelling the namespace?

Injection happens when a pod is created, through a mutating webhook. Pods that existed before the label are untouched. Restart them, for example with kubectl rollout restart deployment -n <ns>, and the new pods come up with the proxy, 2/2. istioctl analyze reports pods without a sidecar as IST0103.

What do the ports 15001, 15006 and 15021 do?

istio-init (or the Istio CNI) sets iptables rules so outbound traffic from the app goes to the proxy on 15001 and inbound traffic to 15006. 15021 is the health port, 15020 serves merged metrics, 15000 is Envoy's admin interface on localhost and 15090 its own stats.

What is a native sidecar?

A container in initContainers with restartPolicy: Always. Kubernetes starts it before the app containers and keeps it running, and it stops after the app, which fixes old problems such as Jobs that never finished because the proxy kept running. Istio uses it by default on current Kubernetes versions.

Which profile should I install?

default (istiod plus an ingress gateway) is the usual start for production, minimal only istiod, demo turns on extra features and logging for trying things out, and ambient installs the sidecar-less mode. istioctl profile list shows them, and --set changes single values.

What does SYNCED mean in istioctl proxy-status?

It means the proxy has acknowledged the latest configuration istiod sent for that type (clusters, listeners, endpoints, routes). STALE means istiod sent something the proxy has not acknowledged, which points at a proxy problem or an overloaded control plane.

In an interview Mid

You labelled a namespace for Istio injection but traffic is not going through the mesh. Why?

Injection happens when a pod is created: a mutating webhook adds the proxy to new pods in a namespace labelled istio-injection=enabled. Pods that already ran stay 1/1, so I check kubectl get pods for 2/2 and run istioctl analyze, which reports IST0103 for pods without a sidecar. The fix is kubectl rollout restart on the Deployments. Then istioctl proxy-status should list every new pod as synced with istiod. If the restarted pods fail to be created at all, the webhook is failing closed, usually because istiod is down.

Also asked: How does Istio redirect a pod's traffic through the sidecar? · What happens to new pods if istiod is unavailable? · What does istioctl analyze check?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.