Why look under the hood
One day a Service answers "Connection refused" and another just hangs, and describe svc looks fine for both. To tell those apart you need to know what actually happens to a packet sent to a ClusterIP. It is less magic than it looks: a few kernel rules on every node, rules you can read.
What you need to know already: Services, ClusterIP and EndpointSlices (16.1), TCP and the three ways a connection fails (9.1), ports (9.8), how Docker's published ports use NAT rules in the kernel (11.15), routing (8.11), DaemonSets (15.22), kube-proxy's job in one line (15.7).
Nothing listens on a ClusterIP
10.96.142.147 is not configured on any network interface, no process has a socket on it, and it does not answer ping. ping -c 2 sends two ICMP echo packets (the "are you there?" packet ping uses):
$ k exec -n shop toolbox -- ping -c 2 web
PING web (10.96.142.147): 56 data bytes
--- web ping statistics ---
2 packets transmitted, 0 packets received, 100% packet loss
Yet wget web/ works. How?
The ClusterIP exists only as packet-rewriting rules on every node, written by kube-proxy (a DaemonSet in kube-system, so one copy per node). When a pod sends the first TCP packet (the SYN, 9.1) to 10.96.142.147:80, the kernel of its own node rewrites the destination to one real pod's IP and port, then routes it normally.
That rewrite is DNAT - destination NAT. NAT (network address translation) means the kernel changes addresses in a packet as it passes; destination NAT changes where it is going. You met the same trick in 11.15: Docker's -p 8080:80 is a DNAT rule from the host's port to the container.
So the Service is a rule, not a server. That is why it costs nothing to scale, and why people are confused the first time they capture traffic with tcpdump (9.11) and never see the ClusterIP arrive anywhere.
kube-proxy watches Services and EndpointSlices through the API server and rewrites the rules when either changes. It is not in the data path - packets never pass through the kube-proxy process. If the kube-proxy pod dies, existing rules keep working; they just stop being updated.
iptables in two minutes
The Linux kernel has a packet filter called netfilter, and iptables is the command that reads and writes its rules. The words you need:
table a group of rules for one purpose. "nat" = rules that rewrite addresses;
"filter" = rules that accept, drop or reject packets.
chain a named, ordered list of rules inside a table. The kernel runs some
built-in chains at fixed points (PREROUTING: a packet just arrived;
OUTPUT: a local process sends one). kube-proxy adds its own chains
(KUBE-SERVICES...) and jumps to them from the built-in ones.
rule "if the packet matches X, do target Y". Checked top to bottom.
target the action: another chain to jump to, or DNAT, REJECT, ACCEPT, DROP...
You will read these with iptables -t TABLE -L CHAIN -n (list one chain) and iptables-save (dump everything).
Looking at the rules: get onto a node
The rules live in the node's kernel, so you need a shell on a node. kubectl debug node/NAME starts a special pod on that node that shares the host's network (so it sees the node's rules) and mounts the node's whole filesystem at /host. -it gives you an interactive terminal (as with docker run -it, 11.1); --image picks the image for the pod:
$ k debug node/worker-1 -it --image=busybox:1.36
Creating debugging pod node-debugger-worker-1-mxww6 with container debugger on node worker-1.
If you don't see a command prompt, try pressing enter.
/ # chroot /host
# iptables -t nat -L KUBE-SERVICES -n
Chain KUBE-SERVICES (2 references)
target prot opt source destination
KUBE-SVC-NPX46M4PTMTKRN6Y 6 -- 0.0.0.0/0 10.96.0.1 /* default/kubernetes:https cluster IP */ tcp dpt:443
KUBE-SVC-TCOU7JCQXEZGVUNU 17 -- 0.0.0.0/0 10.96.0.10 /* kube-system/kube-dns:dns cluster IP */ udp dpt:53
KUBE-SVC-ERIFXISQEP7F7OF4 6 -- 0.0.0.0/0 10.96.0.10 /* kube-system/kube-dns:dns-tcp cluster IP */ tcp dpt:53
KUBE-SVC-E3PXEPMWAWIUT6ZV 6 -- 0.0.0.0/0 10.96.142.147 /* shop/web cluster IP */ tcp dpt:80
KUBE-NODEPORTS 0 -- 0.0.0.0/0 0.0.0.0/0 /* kubernetes service nodeports; NOTE: this must be the last rule in this chain */ ADDRTYPE match dst-type LOCAL
chroot /host makes /host the root directory for your shell, so from now on you run the node's own programs instead of the busybox image's (here Ubuntu's iptables 1.8.10). Then iptables -t nat -L KUBE-SERVICES -n: in the nat table (-t nat), list (-L) the chain KUBE-SERVICES, with numbers instead of names (-n: no DNS lookups, and protocols as numbers: 6 = TCP, 17 = UDP).
The columns: target (where a matching packet goes), prot (protocol), source / destination (0.0.0.0/0 = any address), then a comment kube-proxy wrote and the extra matches (tcp dpt:80 = TCP destination port 80).
The debug pod stays behind after you exit - delete it (k delete pod node-debugger-...).
Read it top to bottom: every packet leaving a pod or arriving at the node passes KUBE-SERVICES (hooked from PREROUTING and OUTPUT). One rule per Service port: "destination 10.96.142.147, TCP port 80 -> jump to KUBE-SVC-E3PXEPMWAWIUT6ZV". The chain name is KUBE-SVC- + the first 16 characters of a hash (a fixed fingerprint) of "shop/web" + "tcp" - so it is the same on every cluster: KUBE-SVC-NPX46M4PTMTKRN6Y is default/kubernetes:https everywhere in the world, handy to recognise.
One Service, followed
iptables-save prints every rule as the command that would create it - one line per rule, easy to grep. -t nat limits it to the nat table. In each line: -A CHAIN = the chain it belongs to, -d / -s destination / source address, -p tcp --dport 80 protocol and port, -m comment a note, -j ("jump") the target:
# iptables-save -t nat | grep shop/web
-A KUBE-SERVICES -d 10.96.142.147/32 -p tcp -m comment --comment "shop/web cluster IP" -m tcp --dport 80 -j KUBE-SVC-E3PXEPMWAWIUT6ZV
-A KUBE-SVC-E3PXEPMWAWIUT6ZV ! -s 10.244.0.0/16 -d 10.96.142.147/32 -p tcp -m comment --comment "shop/web cluster IP" -m tcp --dport 80 -j KUBE-MARK-MASQ
-A KUBE-SVC-E3PXEPMWAWIUT6ZV -m comment --comment "shop/web -> 10.244.1.16:8080" -m statistic --mode random --probability 0.33333333349 -j KUBE-SEP-LGD3B4ARKGRIDQDK
-A KUBE-SVC-E3PXEPMWAWIUT6ZV -m comment --comment "shop/web -> 10.244.2.164:8080" -m statistic --mode random --probability 0.50000000000 -j KUBE-SEP-TZ4NB2HX7ZK6QXWB
-A KUBE-SVC-E3PXEPMWAWIUT6ZV -m comment --comment "shop/web -> 10.244.2.201:8080" -j KUBE-SEP-EF6EG2NV2G6JGAVH
-A KUBE-SEP-LGD3B4ARKGRIDQDK -s 10.244.1.16/32 -m comment --comment shop/web -j KUBE-MARK-MASQ
-A KUBE-SEP-LGD3B4ARKGRIDQDK -p tcp -m comment --comment shop/web -m tcp -j DNAT --to-destination 10.244.1.16:8080
...
Line by line:
KUBE-SERVICESmatches the ClusterIP and port and jumps to the Service chain.- Traffic from outside the pod network (
! -s 10.244.0.0/16: source NOT in the pod range, e.g. a process on the node) is marked for masquerade. Masquerade is SNAT (source NAT): the kernel replaces the sender's address with the node's own, so the reply comes back through this node and gets un-rewritten on the way. - Load balancing is probability. Three endpoints: the first rule is taken with probability 1/3; if not, the second with 1/2 of what is left; the last catches everything else. Each ends up with 1/3. iptables stores the number as a 31-bit fraction, which is why 1/3 prints as 0.33333333349. Four endpoints: 0.25, 0.33333333349, 0.5, rest.
- Each
KUBE-SEP-(service endpoint) chain DNATs to one pod IP:port (--to-destination 10.244.1.16:8080). The masquerade rule above it is the hairpin case: a pod that reaches itself through its own Service.
The decision is made once per connection: the kernel's connection tracker (conntrack, below) remembers which pod this connection went to. So a long-lived connection sticks to one pod. With HTTP/2 or gRPC (protocols that send many requests over one long connection) "load balancing does not work" unless the client spreads its own connections.
The Service with no endpoints
A Service whose selector matches no Ready pod has no KUBE-SVC chain. Instead, a rule in the filter table rejects it. REJECT throws the packet away and tells the sender so (here with an ICMP "port unreachable" message) - unlike DROP, which throws it away silently:
# iptables -t filter -L KUBE-SERVICES -n
Chain KUBE-SERVICES (2 references)
target prot opt source destination
REJECT 6 -- 0.0.0.0/0 10.96.201.7 /* shop/legacy has no endpoints */ tcp dpt:80 reject-with icmp-port-unreachable
So the client sees an instant Connection refused, not a timeout. That difference is diagnostic gold:
Connection refused, instantly no endpoints (REJECT), or endpoints exist but nothing
listens on the targetPort in the pod
timed out the port is not a Service port (no rule matches), or a
NetworkPolicy drops it, or the pod's node is unreachable
(A NetworkPolicy is a firewall rule for pods - 16.29.)
conntrack: the NAT, remembered
conntrack is the kernel's connection-tracking table: for every connection it keeps the original addresses and the rewritten ones, so every later packet of that connection (and every reply) is rewritten the same way without re-running the rules. The conntrack command reads it; -L lists entries, -d 10.96.142.147 only those going to that address:
# conntrack -L -d 10.96.142.147
tcp 6 117 TIME_WAIT src=10.244.2.127 dst=10.96.142.147 sport=51234 dport=80 src=10.244.1.16 dst=10.244.2.127 sport=8080 dport=51234 [ASSURED] mark=0 use=1
conntrack v1.4.8 (conntrack-tools): 1 flow entries have been shown.
One line, one connection. tcp 6 protocol, 117 seconds until the entry expires, TIME_WAIT the TCP state (9.5). The first src= dst= sport= dport= group is the original direction (toolbox -> ClusterIP:80); the second group is the expected reply (pod 10.244.1.16:8080 -> toolbox) - so you can read which pod this connection went to. [ASSURED] = traffic seen both ways. conntrack entries are recorded on the node where the connection started.
The other modes
iptables is only one way kube-proxy can program the kernel. Its mode is set in its ConfigMap (15.29): kubectl get cm kube-proxy -n kube-system -o yaml (cm = configmap), key config.conf. mode: "" means the default: iptables on Linux.
IPVS (IP Virtual Server) is a load balancer built into the Linux kernel: you declare a virtual server (an IP:port) and a list of real servers behind it, and the kernel spreads connections using a scheduler (rr = round robin, one after the other; lc = least connections; sh = source hash, same client -> same server). nftables is the newer replacement for iptables in the kernel, with fast lookup tables instead of long rule lists.
iptables the default. One chain per Service and per endpoint; rule updates
rewrite the whole table (slow with tens of thousands of endpoints).
ipvs kernel load balancer: a virtual server per ClusterIP:port, real
servers = endpoints, schedulers (rr, lc, sh). Every ClusterIP is bound
to a dummy interface kube-ipvs0 - so in IPVS mode ClusterIPs DO answer
ping. Deprecated in Kubernetes 1.35 (warns); off by default in 1.40,
removed in 1.43.
nftables GA since 1.33, the successor to both: set/map lookups instead of long
chains. `nft list table ip kube-proxy` instead of iptables.
In IPVS mode the node's view looks like this. ipvsadm -Ln lists (-L) the virtual servers with numeric addresses (-n):
# ipvsadm -Ln
IP Virtual Server version 1.2.1 (size=4096)
Prot LocalAddress:Port Scheduler Flags
-> RemoteAddress:Port Forward Weight ActiveConn InActConn
TCP 10.96.142.147:80 rr
-> 10.244.1.16:8080 Masq 1 0 0
-> 10.244.2.164:8080 Masq 1 0 0
-> 10.244.2.201:8080 Masq 1 0 0
TCP 10.96.142.147:80 rr is the virtual server (the ClusterIP) with its scheduler; each -> line is a real server (a pod). Forward Masq = reach it by NAT, Weight = its share, ActiveConn / InActConn = open and closing connections.
Changing the mode: edit the ConfigMap, then restart the DaemonSet's pods - kube-proxy reads its config only at start. Line by line: save the ConfigMap to a file; sed -i edits the file in place (7.6) and k apply -f sends it back; rollout restart ds replaces every kube-proxy pod (ds = daemonset); logs -l k8s-app=kube-proxy prints the logs of all pods with that label:
k get cm kube-proxy -n kube-system -o yaml > kp.yaml
sed -i 's/mode: ""/mode: "ipvs"/' kp.yaml && k apply -f kp.yaml
k rollout restart ds kube-proxy -n kube-system
k logs -n kube-system -l k8s-app=kube-proxy | grep Proxier # "Using ipvs Proxier"
A typo in the mode ("ipvss") crash-loops every kube-proxy pod: the rules already programmed keep working, new Services never appear.
Some clusters have no kube-proxy at all: their network plugin (16.35) does the same job with eBPF - small programs loaded into the kernel that handle packets directly. The concepts are the same, the tools to inspect them differ.
What you can now do:
- explain why a ClusterIP works although nothing listens on it
- read the KUBE-SERVICES / KUBE-SVC / KUBE-SEP chains on a node
- tell "refused" (no endpoints) from "timed out" (no matching rule)