OnCallReady

Lesson 16.35 · Kubernetes: Networking & Storage · 14 min read

CNI: how a pod gets its IP, and why the CIDRs matter

In plain words

Think of a new tenant moving into an apartment block. Before they unpack, the building manager (the kubelet) calls the facilities company (the CNI plugin). Facilities runs a cable from the tenant's flat to the building's switchboard, gives the flat an extension number from the building's own range, and updates the building directory. The city's directory already knows "all numbers starting 10.244.1 are in building 1", so calls from other buildings find it.

That's CNI: the kubelet asks containerd for a sandbox, containerd runs the plugins in /etc/cni/net.d, Calico creates a veth pair (eth0 in the pod, caliXXXX on the node), IPAM hands out an IP from the node's pod CIDR, and routes tell every node which node owns which block.

Why know how pods get their IPs

Every pod has its own IP, and any pod can reach any other pod on any node without NAT. Nothing in Linux does that by itself: a network plugin builds it when each pod starts. When pods are stuck in ContainerCreating, or "pod on node A cannot reach pod on node B", or a pod cannot reach a company server because the address ranges collide, you need to know what that plugin did.

What you need to know already: network namespaces, veth pairs and bridges from Docker (11.15), CIDR and overlapping ranges (8.3, 8.6), the routing table (8.11), ARP (8.14), MTU and black holes (9.13), the kubelet and the container runtime (15.7), kubectl debug node and chroot /host (16.3), the service CIDR (16.1).

The CNI plugin

CNI (Container Network Interface) is a small standard: "to connect a container, run this program with these arguments". A CNI plugin is such a program. The main one for a cluster (Calico here; Cilium and Flannel are other common ones) gives every pod an interface and an IP, and makes pod IPs routable between nodes. It is also what enforces NetworkPolicy (16.29).

The sequence

When the kubelet starts a pod:

  1. It asks the container runtime (containerd, over CRI, 15.7) for a pod sandbox: a network namespace (11.15) held open by a tiny pause container, which all the pod's containers then share.
  2. containerd runs the CNI plugins listed in /etc/cni/net.d/*.conflist (the first file alphabetically), programs found in /opt/cni/bin, passing the path of that network namespace.
  3. The plugin (here Calico) creates a veth pair (11.15): one end becomes the pod's eth0, the other stays on the node as caliXXXXXXXXXXX. It asks its IPAM (IP address management: the part that hands out free addresses and remembers who has which) for an address, and installs a route on the node pointing that IP at the veth.
  4. The result goes back to the kubelet, which writes status.podIP.

If step 2 or 3 fails, the pod never gets an IP and sits in ContainerCreating with the event FailedCreatePodSandBox.

Later (Ch 18): a whole-cluster CNI outage, and how to recover from it.

Look at the config on a node:

$ k debug node/worker-1 -it --image=busybox:1.36
/ # chroot /host
# cat /etc/cni/net.d/10-calico.conflist
{
  "name": "k8s-pod-network",
  "cniVersion": "0.3.1",
  "plugins": [
    {
      "type": "calico",
      "datastore_type": "kubernetes",
      "nodename": "worker-1",
      "ipam": {
          "type": "calico-ipam"
      },
      "policy": {
          "type": "k8s"
      },
      ...
    },
    {
      "type": "portmap",
      "snat": true,
      "capabilities": {"portMappings": true}
    },
    ...
# ls /opt/cni/bin
bandwidth  calico  calico-ipam  flannel  host-local  install  loopback  portmap  tuning

It is JSON (7.11). plugins is a chain, run in order: calico does the interface and the IP (with calico-ipam for the address), portmap implements hostPort (a pod port opened on its node, like Docker's -p), bandwidth implements traffic-speed limits.

The pod's side

ip addr lists the pod's interfaces and addresses; ip route its routes:

$ k exec -n shop toolbox -- ip addr
3: eth0@if27: <BROADCAST,MULTICAST,UP,LOWER_UP,M-DOWN> mtu 1480 qdisc noqueue state UP qlen 1000
    inet 10.244.2.127/32 scope global eth0
$ k exec -n shop toolbox -- ip route
default via 169.254.1.1 dev eth0
169.254.1.1 dev eth0 scope link

Two odd things:

The node's side

# ip route
default via 10.64.0.1 dev enp0s1 proto dhcp src 10.64.0.11 metric 100
10.244.0.0/24 via 10.64.0.10 dev tunl0 proto bird onlink
blackhole 10.244.1.0/24 proto bird
10.244.1.16 dev cali10d8a403733 scope link
10.244.1.118 dev cali6323199eb55 scope link
10.244.2.0/24 via 10.64.0.12 dev tunl0 proto bird onlink
10.64.0.0/24 dev enp0s1 proto kernel scope link src 10.64.0.11 metric 100

Read it as a map of the cluster:

That is the whole trick of the pod network: every node knows which node owns which pod block.

Other plugins do the same in other ways. An overlay network wraps pod packets inside node-to-node packets, as IPIP does; VXLAN is another wrapping format (UDP port 4789), used by Flannel and by Calico's VXLAN mode. Without wrapping, plain routes work when the network between nodes knows the pod blocks (Calico BGP to the routers, a cloud's route tables). Cilium uses eBPF (16.3). Some cloud plugins give pods addresses straight from the cloud network, next to the nodes.

The three ranges

-o custom-columns=NAME:path,... (15.38) prints the columns you choose:

$ k get nodes -o custom-columns=NAME:.metadata.name,PODCIDR:.spec.podCIDR,IP:.status.addresses[0].address
NAME       PODCIDR         IP
cp-1       10.244.0.0/24   10.64.0.10
worker-1   10.244.1.0/24   10.64.0.11
worker-2   10.244.2.0/24   10.64.0.12

Each node's podCIDR is its slice of the pod range. Three ranges in every cluster:

node network      10.64.0.0/24      the LAN / cloud subnet the nodes sit in
pod network       10.244.0.0/16     kubeadm init --pod-network-cidr; one /24 per node
service network   10.96.0.0/12      kube-apiserver --service-cluster-ip-range

The pod CIDR is chosen when the cluster is created (kubeadm init, the command that builds a kubeadm cluster); the service CIDR (16.1) is a flag of the API server.

They must not overlap with each other, nor with anything the pods need to reach (8.6). If the pod CIDR overlapped a company network 10.244.0.0/16, a pod calling a server at 10.244.7.9 would route to a local pod block (or a blackhole) instead. The service range is not routed at all - it only exists in kube-proxy's rules (16.3).

Both are hard to change after install (in practice: a new cluster), which is why the IP plan is decided before the cluster is created. Where pods take addresses from the cloud network itself, every pod uses up a subnet address - the subnet-sizing trap of 8.9.

The pod CIDR size also caps pods per node: a /24 per node is 254 addresses; the kubelet's default maxPods is 110.

Which CNI, and does it do policies?

Calico     routes/BGP, IPIP or VXLAN; NetworkPolicy + its own GlobalNetworkPolicy
Cilium     eBPF, can replace kube-proxy; NetworkPolicy + HTTP-level and DNS-name
           policies, and a flow viewer (Hubble)
Flannel    simple VXLAN overlay; NO NetworkPolicy enforcement
cloud CNIs cloud-network IPs (or an overlay); policies via an add-on (often Calico or Cilium)

What you can now do:

Why it helps

When a CNI breaks, pods sit in ContainerCreating with FailedCreatePodSandBox, and the fix lives on the node, not in kubectl. Knowing the sequence (sandbox, conflist, binaries in /opt/cni/bin, IPAM) tells you where to look.

The CIDR lesson matters even more for the job. The pod, service and node ranges must not overlap each other or anything pods need to reach, and they are practically impossible to change after install. At a bank with a large 10.0.0.0/8 corporate network, a pod CIDR that overlaps an on-prem range means pods silently can't reach that server. In a cloud, the choice between pods taking cloud-network IPs and an overlay decides your subnet sizing. You'll be in those design reviews.

FAQ

What is the pause container for?

It holds the pod's network namespace (and other shared namespaces) open. The pod sandbox is created first with the pause container, the CNI plugins configure networking inside that namespace, and then the app containers join it. That's why all containers in a pod share one IP and can talk over localhost, and why an app container can restart without losing the pod IP.

Why is my pod stuck in ContainerCreating with FailedCreatePodSandBox?

The sandbox or its network couldn't be set up, so the pod never got an IP. Typical causes: the CNI agent (calico-node) isn't running on that node, /etc/cni/net.d is empty or has a broken conflist, the plugin binary is missing from /opt/cni/bin, or IPAM ran out of addresses. k describe pod shows the plugin's error; the node's containerd and kubelet logs show more.

Why does the pod have a /32 and a gateway of 169.254.1.1?

That's Calico's design. The pod gets a single address, and its default route points to a link-local gateway that doesn't exist. Calico answers ARP for 169.254.1.1 from the node's side of the veth (proxy ARP), so every packet goes straight into the node's routing table, where Calico's routes decide where it goes. Other CNIs use a bridge and a real gateway instead.

Can I change the pod CIDR after the cluster is built?

In practice, no. The pod CIDR is baked into kubeadm's config, each node's spec.podCIDR, the CNI's IP pools and existing pod addresses; the service range is in the API server flags and every ClusterIP. Changing them means a migration that is riskier than building a new cluster. That's why the IP plan is decided before kubeadm init or az aks create.

Does every CNI support NetworkPolicy?

No. Calico and Cilium enforce NetworkPolicy (and add their own policy CRDs); Cilium adds L7 and FQDN rules and Hubble for visibility. Flannel is a simple VXLAN overlay and enforces nothing, silently. Some cloud CNIs rely on a separate engine (often Calico or Cilium) picked at cluster creation. Check before you trust a policy.

In an interview Mid

What is the difference between the pod CIDR and the service CIDR, and how does a pod get its IP?

Three ranges in every cluster, which must not overlap each other or anything the pods need to reach:

How a pod gets its IP: the kubelet asks containerd (CRI) for a pod sandbox (a network namespace held by the pause container); containerd runs the CNI plugin from /etc/cni/net.d; the plugin (Calico) creates a veth pair, takes an address from its IPAM, adds a node route to it; the kubelet writes status.podIP. Other nodes reach that pod block through a route (here IPIP via tunl0, learned over BGP).

If that fails: ContainerCreating and FailedCreatePodSandBox. Both ranges are hard to change later, so plan them first.

Also asked: What happens, network-wise, when a pod is created? · How does a pod on one node reach a pod on another node? · Why is the pod MTU smaller than the node's?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.