OnCallReady

Lesson 16.3 · Kubernetes: Networking & Storage · 25 min read

How a ClusterIP really routes: kube-proxy, iptables, IPVS

In plain words

Imagine a letter addressed to "The Mayor". There is no house called The Mayor. Instead, every post office in town has a sticker rule: "anything to The Mayor, cross out the address and write one of these three real houses instead". The letter never goes to a Mayor's house; it is re-addressed at the first post office it passes.

A ClusterIP works the same way. Nothing listens on 10.96.142.147. kube-proxy writes rewrite rules (iptables, IPVS or nftables) into every node's kernel, and the node where the packet starts DNATs it to one pod IP:port, chosen at random per connection. conntrack remembers the choice so the rest of the conversation goes to the same pod.

Why look under the hood

One day a Service answers "Connection refused" and another just hangs, and describe svc looks fine for both. To tell those apart you need to know what actually happens to a packet sent to a ClusterIP. It is less magic than it looks: a few kernel rules on every node, rules you can read.

What you need to know already: Services, ClusterIP and EndpointSlices (16.1), TCP and the three ways a connection fails (9.1), ports (9.8), how Docker's published ports use NAT rules in the kernel (11.15), routing (8.11), DaemonSets (15.22), kube-proxy's job in one line (15.7).

Nothing listens on a ClusterIP

10.96.142.147 is not configured on any network interface, no process has a socket on it, and it does not answer ping. ping -c 2 sends two ICMP echo packets (the "are you there?" packet ping uses):

$ k exec -n shop toolbox -- ping -c 2 web
PING web (10.96.142.147): 56 data bytes

--- web ping statistics ---
2 packets transmitted, 0 packets received, 100% packet loss

Yet wget web/ works. How?

The ClusterIP exists only as packet-rewriting rules on every node, written by kube-proxy (a DaemonSet in kube-system, so one copy per node). When a pod sends the first TCP packet (the SYN, 9.1) to 10.96.142.147:80, the kernel of its own node rewrites the destination to one real pod's IP and port, then routes it normally.

That rewrite is DNAT - destination NAT. NAT (network address translation) means the kernel changes addresses in a packet as it passes; destination NAT changes where it is going. You met the same trick in 11.15: Docker's -p 8080:80 is a DNAT rule from the host's port to the container.

So the Service is a rule, not a server. That is why it costs nothing to scale, and why people are confused the first time they capture traffic with tcpdump (9.11) and never see the ClusterIP arrive anywhere.

kube-proxy watches Services and EndpointSlices through the API server and rewrites the rules when either changes. It is not in the data path - packets never pass through the kube-proxy process. If the kube-proxy pod dies, existing rules keep working; they just stop being updated.

iptables in two minutes

The Linux kernel has a packet filter called netfilter, and iptables is the command that reads and writes its rules. The words you need:

table    a group of rules for one purpose. "nat" = rules that rewrite addresses;
         "filter" = rules that accept, drop or reject packets.
chain    a named, ordered list of rules inside a table. The kernel runs some
         built-in chains at fixed points (PREROUTING: a packet just arrived;
         OUTPUT: a local process sends one). kube-proxy adds its own chains
         (KUBE-SERVICES...) and jumps to them from the built-in ones.
rule     "if the packet matches X, do target Y". Checked top to bottom.
target   the action: another chain to jump to, or DNAT, REJECT, ACCEPT, DROP...

You will read these with iptables -t TABLE -L CHAIN -n (list one chain) and iptables-save (dump everything).

Looking at the rules: get onto a node

The rules live in the node's kernel, so you need a shell on a node. kubectl debug node/NAME starts a special pod on that node that shares the host's network (so it sees the node's rules) and mounts the node's whole filesystem at /host. -it gives you an interactive terminal (as with docker run -it, 11.1); --image picks the image for the pod:

$ k debug node/worker-1 -it --image=busybox:1.36
Creating debugging pod node-debugger-worker-1-mxww6 with container debugger on node worker-1.
If you don't see a command prompt, try pressing enter.
/ # chroot /host
# iptables -t nat -L KUBE-SERVICES -n
Chain KUBE-SERVICES (2 references)
target     prot opt source               destination
KUBE-SVC-NPX46M4PTMTKRN6Y  6    --  0.0.0.0/0            10.96.0.1            /* default/kubernetes:https cluster IP */ tcp dpt:443
KUBE-SVC-TCOU7JCQXEZGVUNU  17   --  0.0.0.0/0            10.96.0.10           /* kube-system/kube-dns:dns cluster IP */ udp dpt:53
KUBE-SVC-ERIFXISQEP7F7OF4  6    --  0.0.0.0/0            10.96.0.10           /* kube-system/kube-dns:dns-tcp cluster IP */ tcp dpt:53
KUBE-SVC-E3PXEPMWAWIUT6ZV  6    --  0.0.0.0/0            10.96.142.147        /* shop/web cluster IP */ tcp dpt:80
KUBE-NODEPORTS  0    --  0.0.0.0/0            0.0.0.0/0            /* kubernetes service nodeports; NOTE: this must be the last rule in this chain */ ADDRTYPE match dst-type LOCAL

chroot /host makes /host the root directory for your shell, so from now on you run the node's own programs instead of the busybox image's (here Ubuntu's iptables 1.8.10). Then iptables -t nat -L KUBE-SERVICES -n: in the nat table (-t nat), list (-L) the chain KUBE-SERVICES, with numbers instead of names (-n: no DNS lookups, and protocols as numbers: 6 = TCP, 17 = UDP).

The columns: target (where a matching packet goes), prot (protocol), source / destination (0.0.0.0/0 = any address), then a comment kube-proxy wrote and the extra matches (tcp dpt:80 = TCP destination port 80).

The debug pod stays behind after you exit - delete it (k delete pod node-debugger-...).

Read it top to bottom: every packet leaving a pod or arriving at the node passes KUBE-SERVICES (hooked from PREROUTING and OUTPUT). One rule per Service port: "destination 10.96.142.147, TCP port 80 -> jump to KUBE-SVC-E3PXEPMWAWIUT6ZV". The chain name is KUBE-SVC- + the first 16 characters of a hash (a fixed fingerprint) of "shop/web" + "tcp" - so it is the same on every cluster: KUBE-SVC-NPX46M4PTMTKRN6Y is default/kubernetes:https everywhere in the world, handy to recognise.

One Service, followed

iptables-save prints every rule as the command that would create it - one line per rule, easy to grep. -t nat limits it to the nat table. In each line: -A CHAIN = the chain it belongs to, -d / -s destination / source address, -p tcp --dport 80 protocol and port, -m comment a note, -j ("jump") the target:

# iptables-save -t nat | grep shop/web
-A KUBE-SERVICES -d 10.96.142.147/32 -p tcp -m comment --comment "shop/web cluster IP" -m tcp --dport 80 -j KUBE-SVC-E3PXEPMWAWIUT6ZV
-A KUBE-SVC-E3PXEPMWAWIUT6ZV ! -s 10.244.0.0/16 -d 10.96.142.147/32 -p tcp -m comment --comment "shop/web cluster IP" -m tcp --dport 80 -j KUBE-MARK-MASQ
-A KUBE-SVC-E3PXEPMWAWIUT6ZV -m comment --comment "shop/web -> 10.244.1.16:8080" -m statistic --mode random --probability 0.33333333349 -j KUBE-SEP-LGD3B4ARKGRIDQDK
-A KUBE-SVC-E3PXEPMWAWIUT6ZV -m comment --comment "shop/web -> 10.244.2.164:8080" -m statistic --mode random --probability 0.50000000000 -j KUBE-SEP-TZ4NB2HX7ZK6QXWB
-A KUBE-SVC-E3PXEPMWAWIUT6ZV -m comment --comment "shop/web -> 10.244.2.201:8080" -j KUBE-SEP-EF6EG2NV2G6JGAVH
-A KUBE-SEP-LGD3B4ARKGRIDQDK -s 10.244.1.16/32 -m comment --comment shop/web -j KUBE-MARK-MASQ
-A KUBE-SEP-LGD3B4ARKGRIDQDK -p tcp -m comment --comment shop/web -m tcp -j DNAT --to-destination 10.244.1.16:8080
...

Line by line:

  1. KUBE-SERVICES matches the ClusterIP and port and jumps to the Service chain.
  2. Traffic from outside the pod network (! -s 10.244.0.0/16: source NOT in the pod range, e.g. a process on the node) is marked for masquerade. Masquerade is SNAT (source NAT): the kernel replaces the sender's address with the node's own, so the reply comes back through this node and gets un-rewritten on the way.
  3. Load balancing is probability. Three endpoints: the first rule is taken with probability 1/3; if not, the second with 1/2 of what is left; the last catches everything else. Each ends up with 1/3. iptables stores the number as a 31-bit fraction, which is why 1/3 prints as 0.33333333349. Four endpoints: 0.25, 0.33333333349, 0.5, rest.
  4. Each KUBE-SEP- (service endpoint) chain DNATs to one pod IP:port (--to-destination 10.244.1.16:8080). The masquerade rule above it is the hairpin case: a pod that reaches itself through its own Service.

The decision is made once per connection: the kernel's connection tracker (conntrack, below) remembers which pod this connection went to. So a long-lived connection sticks to one pod. With HTTP/2 or gRPC (protocols that send many requests over one long connection) "load balancing does not work" unless the client spreads its own connections.

The Service with no endpoints

A Service whose selector matches no Ready pod has no KUBE-SVC chain. Instead, a rule in the filter table rejects it. REJECT throws the packet away and tells the sender so (here with an ICMP "port unreachable" message) - unlike DROP, which throws it away silently:

# iptables -t filter -L KUBE-SERVICES -n
Chain KUBE-SERVICES (2 references)
target     prot opt source               destination
REJECT     6    --  0.0.0.0/0            10.96.201.7          /* shop/legacy has no endpoints */ tcp dpt:80 reject-with icmp-port-unreachable

So the client sees an instant Connection refused, not a timeout. That difference is diagnostic gold:

Connection refused, instantly   no endpoints (REJECT), or endpoints exist but nothing
                                listens on the targetPort in the pod
timed out                       the port is not a Service port (no rule matches), or a
                                NetworkPolicy drops it, or the pod's node is unreachable

(A NetworkPolicy is a firewall rule for pods - 16.29.)

conntrack: the NAT, remembered

conntrack is the kernel's connection-tracking table: for every connection it keeps the original addresses and the rewritten ones, so every later packet of that connection (and every reply) is rewritten the same way without re-running the rules. The conntrack command reads it; -L lists entries, -d 10.96.142.147 only those going to that address:

# conntrack -L -d 10.96.142.147
tcp      6 117 TIME_WAIT src=10.244.2.127 dst=10.96.142.147 sport=51234 dport=80 src=10.244.1.16 dst=10.244.2.127 sport=8080 dport=51234 [ASSURED] mark=0 use=1
conntrack v1.4.8 (conntrack-tools): 1 flow entries have been shown.

One line, one connection. tcp 6 protocol, 117 seconds until the entry expires, TIME_WAIT the TCP state (9.5). The first src= dst= sport= dport= group is the original direction (toolbox -> ClusterIP:80); the second group is the expected reply (pod 10.244.1.16:8080 -> toolbox) - so you can read which pod this connection went to. [ASSURED] = traffic seen both ways. conntrack entries are recorded on the node where the connection started.

The other modes

iptables is only one way kube-proxy can program the kernel. Its mode is set in its ConfigMap (15.29): kubectl get cm kube-proxy -n kube-system -o yaml (cm = configmap), key config.conf. mode: "" means the default: iptables on Linux.

IPVS (IP Virtual Server) is a load balancer built into the Linux kernel: you declare a virtual server (an IP:port) and a list of real servers behind it, and the kernel spreads connections using a scheduler (rr = round robin, one after the other; lc = least connections; sh = source hash, same client -> same server). nftables is the newer replacement for iptables in the kernel, with fast lookup tables instead of long rule lists.

iptables   the default. One chain per Service and per endpoint; rule updates
           rewrite the whole table (slow with tens of thousands of endpoints).
ipvs       kernel load balancer: a virtual server per ClusterIP:port, real
           servers = endpoints, schedulers (rr, lc, sh). Every ClusterIP is bound
           to a dummy interface kube-ipvs0 - so in IPVS mode ClusterIPs DO answer
           ping. Deprecated in Kubernetes 1.35 (warns); off by default in 1.40,
           removed in 1.43.
nftables   GA since 1.33, the successor to both: set/map lookups instead of long
           chains. `nft list table ip kube-proxy` instead of iptables.

In IPVS mode the node's view looks like this. ipvsadm -Ln lists (-L) the virtual servers with numeric addresses (-n):

# ipvsadm -Ln
IP Virtual Server version 1.2.1 (size=4096)
Prot LocalAddress:Port Scheduler Flags
  -> RemoteAddress:Port           Forward Weight ActiveConn InActConn
TCP  10.96.142.147:80 rr
  -> 10.244.1.16:8080             Masq    1      0          0
  -> 10.244.2.164:8080            Masq    1      0          0
  -> 10.244.2.201:8080            Masq    1      0          0

TCP 10.96.142.147:80 rr is the virtual server (the ClusterIP) with its scheduler; each -> line is a real server (a pod). Forward Masq = reach it by NAT, Weight = its share, ActiveConn / InActConn = open and closing connections.

Changing the mode: edit the ConfigMap, then restart the DaemonSet's pods - kube-proxy reads its config only at start. Line by line: save the ConfigMap to a file; sed -i edits the file in place (7.6) and k apply -f sends it back; rollout restart ds replaces every kube-proxy pod (ds = daemonset); logs -l k8s-app=kube-proxy prints the logs of all pods with that label:

k get cm kube-proxy -n kube-system -o yaml > kp.yaml
sed -i 's/mode: ""/mode: "ipvs"/' kp.yaml && k apply -f kp.yaml
k rollout restart ds kube-proxy -n kube-system
k logs -n kube-system -l k8s-app=kube-proxy | grep Proxier    # "Using ipvs Proxier"

A typo in the mode ("ipvss") crash-loops every kube-proxy pod: the rules already programmed keep working, new Services never appear.

Some clusters have no kube-proxy at all: their network plugin (16.35) does the same job with eBPF - small programs loaded into the kernel that handle packets directly. The concepts are the same, the tools to inspect them differ.

What you can now do:

Why it helps

This is the lesson that turns "the network is weird" into a diagnosis. When ping web fails but wget web works, you'll know it's normal. When a client gets Connection refused instantly, you'll know it's the REJECT rule for a Service with no endpoints; a timeout means no rule matched or something dropped it. That split alone saves an hour in an incident channel.

It explains production behaviour people blame on apps: gRPC and HTTP/2 connections sticking to one pod, uneven load after a scale-up (old connections stay put), and kube-proxy crash-looping after a bad ConfigMap edit while existing Services keep working. And when a platform team debates iptables vs IPVS vs nftables vs Cilium's eBPF replacement, you'll know what each actually changes.

FAQ

Why can't I ping a ClusterIP?

Because in iptables (and nftables) mode the ClusterIP is not configured on any interface; it only exists as NAT rules for specific TCP/UDP ports. ICMP echo matches no rule, so nothing answers. In IPVS mode every ClusterIP is bound to the dummy interface kube-ipvs0, so ping does answer there. Either way, ping tells you nothing about whether the Service works; test the port.

If kube-proxy dies, does traffic stop?

No. kube-proxy is not in the data path; it only programs the kernel. Existing rules keep forwarding traffic. What stops is updating: new Services never appear, and endpoints of rescheduled pods are not added, so over time traffic goes to IPs that no longer exist. That is why a crash-looping kube-proxy (for example after a typo in its mode) can go unnoticed for a while.

Why do the probabilities look like 0.333 then 0.5 then nothing?

Rules are evaluated in order. With three endpoints, the first is taken with probability 1/3. If not taken, the second is taken with 1/2 of the remaining 2/3, which is 1/3 overall. The last rule catches the rest, also 1/3. iptables stores the fraction in 31 bits, so 1/3 prints as 0.33333333349.

Which mode should a cluster use: iptables, IPVS or nftables?

iptables is the Linux default and fine for most clusters; its weakness is that updates rewrite the whole table, which gets slow with tens of thousands of endpoints. IPVS scaled better but is deprecated as of Kubernetes 1.35. nftables is GA since 1.33 and is the successor, using set lookups instead of long chains. Many managed and Cilium clusters replace kube-proxy with eBPF entirely.

Where do I look at the rules on a node?

Get a shell on the node (k debug node/worker-1 -it --image=busybox:1.36 then chroot /host, or SSH), then iptables-save -t nat | grep shop/web for one Service, iptables -t nat -L KUBE-SERVICES -n for the entry chain, ipvsadm -Ln in IPVS mode, nft list table ip kube-proxy in nftables mode, and conntrack -L -d <ClusterIP> for live connections. Delete the debug pod afterwards.

In an interview Mid

How does kube-proxy implement a ClusterIP Service?

Nothing listens on a ClusterIP; it exists only as packet-rewriting rules on every node, written by kube-proxy (a DaemonSet) from Services and EndpointSlices. In the default iptables mode:

  1. KUBE-SERVICES (in the nat table) matches ClusterIP + port and jumps to the Service's KUBE-SVC-... chain.
  2. That chain picks an endpoint by probability (1/3, then 1/2, then the rest for three endpoints).
  3. Each KUBE-SEP-... chain DNATs to one pod IP:port; traffic from outside the pod network is masqueraded (SNAT) so replies come back the same way.
  4. conntrack remembers the choice, so the decision is once per connection - a long-lived HTTP/2 or gRPC connection sticks to one pod.

kube-proxy is not in the data path: if it dies, the rules keep working, they just stop updating. A Service with no endpoints gets a REJECT rule instead - an instant "connection refused". Other modes: IPVS (kernel load balancer), nftables, or eBPF plugins that replace kube-proxy.

Also asked: Why does ping to a ClusterIP not work, while a TCP connection does? · Why can gRPC traffic all land on one pod behind a Service? · What does "connection refused" from a Service tell you, compared with a timeout?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.