Why know how pods get their IPs
Every pod has its own IP, and any pod can reach any other pod on any node without NAT. Nothing in Linux does that by itself: a network plugin builds it when each pod starts. When pods are stuck in ContainerCreating, or "pod on node A cannot reach pod on node B", or a pod cannot reach a company server because the address ranges collide, you need to know what that plugin did.
What you need to know already: network namespaces, veth pairs and bridges from Docker (11.15), CIDR and overlapping ranges (8.3, 8.6), the routing table (8.11), ARP (8.14), MTU and black holes (9.13), the kubelet and the container runtime (15.7), kubectl debug node and chroot /host (16.3), the service CIDR (16.1).
The CNI plugin
CNI (Container Network Interface) is a small standard: "to connect a container, run this program with these arguments". A CNI plugin is such a program. The main one for a cluster (Calico here; Cilium and Flannel are other common ones) gives every pod an interface and an IP, and makes pod IPs routable between nodes. It is also what enforces NetworkPolicy (16.29).
The sequence
When the kubelet starts a pod:
- It asks the container runtime (containerd, over CRI, 15.7) for a pod sandbox: a network namespace (11.15) held open by a tiny
pausecontainer, which all the pod's containers then share. - containerd runs the CNI plugins listed in
/etc/cni/net.d/*.conflist(the first file alphabetically), programs found in/opt/cni/bin, passing the path of that network namespace. - The plugin (here Calico) creates a veth pair (11.15): one end becomes the pod's
eth0, the other stays on the node ascaliXXXXXXXXXXX. It asks its IPAM (IP address management: the part that hands out free addresses and remembers who has which) for an address, and installs a route on the node pointing that IP at the veth. - The result goes back to the kubelet, which writes
status.podIP.
If step 2 or 3 fails, the pod never gets an IP and sits in ContainerCreating with the event FailedCreatePodSandBox.
Later (Ch 18): a whole-cluster CNI outage, and how to recover from it.
Look at the config on a node:
$ k debug node/worker-1 -it --image=busybox:1.36
/ # chroot /host
# cat /etc/cni/net.d/10-calico.conflist
{
"name": "k8s-pod-network",
"cniVersion": "0.3.1",
"plugins": [
{
"type": "calico",
"datastore_type": "kubernetes",
"nodename": "worker-1",
"ipam": {
"type": "calico-ipam"
},
"policy": {
"type": "k8s"
},
...
},
{
"type": "portmap",
"snat": true,
"capabilities": {"portMappings": true}
},
...
# ls /opt/cni/bin
bandwidth calico calico-ipam flannel host-local install loopback portmap tuning
It is JSON (7.11). plugins is a chain, run in order: calico does the interface and the IP (with calico-ipam for the address), portmap implements hostPort (a pod port opened on its node, like Docker's -p), bandwidth implements traffic-speed limits.
The pod's side
ip addr lists the pod's interfaces and addresses; ip route its routes:
$ k exec -n shop toolbox -- ip addr
3: eth0@if27: <BROADCAST,MULTICAST,UP,LOWER_UP,M-DOWN> mtu 1480 qdisc noqueue state UP qlen 1000
inet 10.244.2.127/32 scope global eth0
$ k exec -n shop toolbox -- ip route
default via 169.254.1.1 dev eth0
169.254.1.1 dev eth0 scope link
Two odd things:
- The address is a /32 - a "network" of one address - and the gateway 169.254.1.1 does not exist. Calico answers ARP for 169.254.1.1 from the node's end of the veth (proxy ARP: answering ARP on behalf of another address), so every packet leaves the pod straight into the node's routing table.
mtu 1480: 1500 minus the 20-byte header of the tunnel between nodes (below).
The node's side
# ip route
default via 10.64.0.1 dev enp0s1 proto dhcp src 10.64.0.11 metric 100
10.244.0.0/24 via 10.64.0.10 dev tunl0 proto bird onlink
blackhole 10.244.1.0/24 proto bird
10.244.1.16 dev cali10d8a403733 scope link
10.244.1.118 dev cali6323199eb55 scope link
10.244.2.0/24 via 10.64.0.12 dev tunl0 proto bird onlink
10.64.0.0/24 dev enp0s1 proto kernel scope link src 10.64.0.11 metric 100
Read it as a map of the cluster:
- each local pod has a route of its own to its cali interface;
- this node's whole block (
10.244.1.0/24) is a blackhole route (packets are thrown away), so an unused address in it fails fast instead of leaking out; - other nodes' blocks go via
tunl0to that node's IP.tunl0is an IP-in-IP tunnel (IPIP): the pod packet is wrapped inside a second IP packet from node to node - that wrapper is the 20 bytes. proto birdsays who added the route: BIRD, the routing program in Calico's per-node agent (calico-node). The agents tell each other which block lives where over BGP, the protocol routers use to exchange routes.
That is the whole trick of the pod network: every node knows which node owns which pod block.
Other plugins do the same in other ways. An overlay network wraps pod packets inside node-to-node packets, as IPIP does; VXLAN is another wrapping format (UDP port 4789), used by Flannel and by Calico's VXLAN mode. Without wrapping, plain routes work when the network between nodes knows the pod blocks (Calico BGP to the routers, a cloud's route tables). Cilium uses eBPF (16.3). Some cloud plugins give pods addresses straight from the cloud network, next to the nodes.
The three ranges
-o custom-columns=NAME:path,... (15.38) prints the columns you choose:
$ k get nodes -o custom-columns=NAME:.metadata.name,PODCIDR:.spec.podCIDR,IP:.status.addresses[0].address
NAME PODCIDR IP
cp-1 10.244.0.0/24 10.64.0.10
worker-1 10.244.1.0/24 10.64.0.11
worker-2 10.244.2.0/24 10.64.0.12
Each node's podCIDR is its slice of the pod range. Three ranges in every cluster:
node network 10.64.0.0/24 the LAN / cloud subnet the nodes sit in
pod network 10.244.0.0/16 kubeadm init --pod-network-cidr; one /24 per node
service network 10.96.0.0/12 kube-apiserver --service-cluster-ip-range
The pod CIDR is chosen when the cluster is created (kubeadm init, the command that builds a kubeadm cluster); the service CIDR (16.1) is a flag of the API server.
They must not overlap with each other, nor with anything the pods need to reach (8.6). If the pod CIDR overlapped a company network 10.244.0.0/16, a pod calling a server at 10.244.7.9 would route to a local pod block (or a blackhole) instead. The service range is not routed at all - it only exists in kube-proxy's rules (16.3).
Both are hard to change after install (in practice: a new cluster), which is why the IP plan is decided before the cluster is created. Where pods take addresses from the cloud network itself, every pod uses up a subnet address - the subnet-sizing trap of 8.9.
The pod CIDR size also caps pods per node: a /24 per node is 254 addresses; the kubelet's default maxPods is 110.
Which CNI, and does it do policies?
Calico routes/BGP, IPIP or VXLAN; NetworkPolicy + its own GlobalNetworkPolicy
Cilium eBPF, can replace kube-proxy; NetworkPolicy + HTTP-level and DNS-name
policies, and a flow viewer (Hubble)
Flannel simple VXLAN overlay; NO NetworkPolicy enforcement
cloud CNIs cloud-network IPs (or an overlay); policies via an add-on (often Calico or Cilium)
What you can now do:
- explain the steps from "pod scheduled" to "pod has an IP"
- read a node's route table and say which node owns a pod IP
- check that node, pod and service ranges do not overlap, and why that matters