The problem: a ClusterIP only works inside
A ClusterIP answers only inside the cluster: from pods and from the nodes. Your laptop, a customer's browser, a server elsewhere in the company - none of them can use 10.96.142.147. And sometimes you want the opposite of hiding pods behind one IP: a database's replicas need to find each other, one by one. Service types cover these cases.
What you need to know already: Services, ClusterIP, endpoints (16.1), how kube-proxy rewrites packets, DNAT and masquerade/SNAT (16.3), Docker published ports (11.15), load balancers (9.23), ARP (8.14), DNS records A and CNAME (8.22), StatefulSets and their headless Service (15.19), curl (9.21).
The ladder
Each type builds on the one before it:
ClusterIP a virtual IP, reachable only from inside the cluster (pods, nodes)
NodePort ClusterIP + a port 30000-32767 opened on EVERY node's IP
LoadBalancer NodePort + an external load balancer (from a cloud or MetalLB)
that sends traffic to those node ports (or straight to pods)
ExternalName no IP, no proxying: DNS returns a CNAME to another name
headless clusterIP: None - no virtual IP; DNS returns the pod IPs themselves
The type is the spec.type field of the Service (headless is a ClusterIP Service with clusterIP: None). The rest of this lesson takes them one by one.
NodePort
A NodePort Service opens the same port on every node's own IP. Anyone who can reach a node can reach the Service: http://<any-node-ip>:<nodePort>. It is Docker's -p 8080:80 (11.15), but on every node of the cluster at once.
Something to expose - one agnhost pod, pinned to worker-1 so the curls below know where it runs. The patch adds a nodeSelector (run only on nodes with this label; kubernetes.io/hostname is the label that holds a node's name):
$ k delete ns edge --ignore-not-found; k create ns edge
$ k create deployment hello -n edge --image=registry.k8s.io/e2e-test-images/agnhost:2.53 -- /agnhost netexec --http-port=8080
$ k patch deploy hello -n edge -p '{"spec":{"template":{"spec":{"nodeSelector":{"kubernetes.io/hostname":"worker-1"}}}}}'
create service nodeport writes the Service: --tcp=80:8080 is port : targetPort, --node-port=30900 picks the port opened on the nodes:
$ k create service nodeport hello --tcp=80:8080 --node-port=30900 -n edge
service/hello created
$ k get svc hello -n edge
NAME TYPE CLUSTER-IP EXTERNAL-IP PORT(S) AGE
hello NodePort 10.96.173.130 <none> 80:30900/TCP 3s
In PORT(S), 80:30900 = Service port : node port. It still has a ClusterIP: each type adds to the one before. kube-proxy on every node accepts 30900 - including nodes that run none of the pods, and the control plane. From oncall-lab (outside the cluster; the nodes are 10.64.0.10 cp-1, .11 worker-1, .12 worker-2), curl -s (silent, no progress meter) and ; echo for a newline:
$ curl -s http://10.64.0.11:30900/hostname; echo
hello-6c7d9f8b5d-k2x9p
$ curl -s http://10.64.0.10:30900/hostname; echo
hello-6c7d9f8b5d-k2x9p
The pod runs on worker-1 only, yet cp-1 (10.64.0.10) answered too: that node forwarded the connection to worker-1.
Omit nodePort and one is picked for you from the node port range (30000-32767 by default); ask for one outside the range and the API server refuses:
The Service "hello" is invalid: spec.ports[0].nodePort: Invalid value: 80: provided port is not in the valid range. The range of valid ports is 30000-32767
NodePorts are how every external load balancer reaches a cluster underneath, but handing them to users directly is rare: odd port numbers, and you must keep track of node IPs yourself.
externalTrafficPolicy: whose IP does the pod see?
With the default externalTrafficPolicy: Cluster, the node that receives the packet may forward it to a pod on another node. To get the reply back through itself, it SNATs (16.3) the source address to its own:
$ curl -s http://10.64.0.12:30900/clientip; echo
10.244.2.1:43457 <- worker-2's tunnel address, not your 10.64.0.2
/clientip shows the address the pod saw. You came from oncall-lab (10.64.0.2), but the pod sees worker-2's own address on the pod network.
externalTrafficPolicy: Local means "only send outside traffic to pods on the node it arrived at". That changes four things at once:
+ the client source IP is preserved (no SNAT) - audit logs, rate limits and
geo rules need this
+ no extra hop between nodes
- a node with no local endpoint DROPS the traffic (a timeout, not a refusal)
- load follows pod placement, not node count: 3 pods on node A and 1 on node B
behind a balancer that splits 50/50 gives the lone pod half the traffic
k patch svc ... -p '{json}' merges that JSON into the Service (15.40). curl -sS -m 3: silent but still show errors (-S), give up after 3 seconds (-m 3):
$ k patch svc hello -n edge -p '{"spec":{"externalTrafficPolicy":"Local"}}'
$ curl -s http://10.64.0.11:30900/clientip; echo # the node with the pod
10.64.0.2:40211
$ curl -sS -m 3 http://10.64.0.12:30900/clientip # the node without one
curl: (28) Connection timed out after 3002 milliseconds
(internalTrafficPolicy, seen in describe svc, is the same idea for traffic from inside the cluster: Local = only to pods on the caller's own node.)
For LoadBalancer + Local, the API server also picks a healthCheckNodePort: a port the external balancer checks on each node, so it only sends traffic to nodes that currently have an endpoint.
LoadBalancer
A LoadBalancer Service asks for a real load balancer (9.23) with its own address in front of the node ports. Kubernetes does not build that box itself. In a cloud, a control-plane component called the cloud-controller-manager (the part that talks to the cloud provider's API) sees type: LoadBalancer, creates a load balancer at the cloud provider, and writes its address into the Service's status.loadBalancer.ingress - which get svc shows as EXTERNAL-IP.
On a bare kubeadm cluster like this one there is nobody to do that:
$ k create service loadbalancer shop-lb -n edge --tcp=80:8080
service/shop-lb created
$ k get svc shop-lb -n edge
NAME TYPE CLUSTER-IP EXTERNAL-IP PORT(S) AGE
shop-lb LoadBalancer 10.96.12.103 <pending> 80:31746/TCP 12m
EXTERNAL-IP <pending> forever, and no events explaining why. Everything below LoadBalancer on the ladder (the ClusterIP, the node port 31746) works already.
MetalLB fills the gap on bare metal (your own machines, no cloud). It has two parts: a controller that hands out IPs from an IPAddressPool (a range of free addresses you give it), and speakers (a DaemonSet, one per node) that make the network deliver those IPs to a node. In layer-2 mode one node's speaker answers ARP (8.14) for the IP - "that IP is at my MAC address" - so traffic for it arrives at that node. An L2Advertisement is the object that turns this on for a pool.
MetalLB's objects are extra kinds added to the API (apiVersion: metallb.io/v1beta1), written like any other manifest:
apiVersion: metallb.io/v1beta1
kind: IPAddressPool
metadata: {name: lab-pool, namespace: metallb-system}
spec:
addresses: [10.64.0.240-10.64.0.250] # free addresses on the node LAN
---
apiVersion: metallb.io/v1beta1
kind: L2Advertisement
metadata: {name: lab-l2, namespace: metallb-system}
spec:
ipAddressPools: [lab-pool]
Once installed, the Service's events (the end of describe, tail -3 = the last three lines) say which IP it got and which node announces it:
# once MetalLB is installed with a pool (the LoadBalancer mission)
k describe svc shop-lb | tail -3
Normal IPAllocated 2s metallb-controller Assigned IP ["10.64.0.240"]
Normal nodeAssigned 2s metallb-speaker announcing from node "worker-2" with protocol "layer2"
The pool without an L2Advertisement is the classic half-install: the Service gets an EXTERNAL-IP, but nobody answers ARP for it, so curl reports No route to host (the ARP failure of 8.14). In a cloud, the equivalent knobs (internal vs public load balancer, and so on) are annotations (15.26) on the Service that the cloud-controller-manager reads.
ExternalName
spec:
type: ExternalName
externalName: db.lab
An ExternalName Service is only a DNS alias: asking for billing-db returns a CNAME (8.22) pointing at db.lab. nslookup shows the chain:
# billing-db = the ExternalName above, in db-lab (the headless mission)
k exec toolbox -- nslookup billing-db
billing-db.db-lab.svc.cluster.local canonical name = db.lab.
Name: db.lab
Address: 10.0.3.12
"canonical name =" is the CNAME; the last lines are the real address.
Only DNS. No ClusterIP, no ports honoured, no kube-proxy rules, and a NetworkPolicy (16.29) cannot target it: the client connects to 10.0.3.12 itself. Useful to give an external database an in-cluster name you can later repoint. Traps: HTTP clients send Host: billing-db (9.21), and TLS clients expect a certificate for billing-db (9.15) - neither matches what the external server expects.
Headless
You met the headless Service in 15.19: clusterIP: None.
spec:
clusterIP: None
selector: {app: ledger}
ports: [{name: pg, port: 5432}]
No virtual IP, no kube-proxy rules. DNS returns an A record (8.22) per ready pod, and with a StatefulSet (serviceName: ledger) each pod gets its own stable name. dig +short prints only the answer (8.18):
# ledger = a StatefulSet + headless Service in db-lab (the headless mission)
k exec t -- nslookup ledger
Name: ledger.db-lab.svc.cluster.local
Address: 10.244.1.16
Name: ledger.db-lab.svc.cluster.local
Address: 10.244.2.127
k exec t -- dig +short ledger-0.ledger.db-lab.svc.cluster.local
10.244.2.127
That is how database replicas find "the primary is ledger-0", and how a client that does its own load balancing (gRPC clients often do) sees every backend.
publishNotReadyAddresses: true publishes pods in DNS before they are Ready: the members of a clustered database need to find each other to become ready at all.
clusterIP is immutable (cannot be changed after creation). Turning a ClusterIP Service into a headless one means delete and recreate:
$ k create service clusterip ledger-ro -n edge --tcp=5432
service/ledger-ro created
$ k patch svc ledger-ro -n edge -p '{"spec":{"clusterIP":"None"}}'
The Service "ledger-ro" is invalid: spec.clusterIPs[0]: Invalid value: []string{"None"}: may not change once set
sessionAffinity
sessionAffinity: ClientIP sends the same client IP to the same pod every time, for sessionAffinityConfig.clientIP.timeoutSeconds (default 10800 = 3h). kube-proxy does it with an extra iptables module (recent) that remembers recent senders.
The catch: behind SNAT (externalTrafficPolicy Cluster, or a company proxy) every user shares one client IP and lands on one pod. It is a crutch for apps that keep login sessions in memory; the fix is to not do that.
Choosing
only other pods call it ClusterIP
a StatefulSet, peers find each other, gRPC LB headless (+ a ClusterIP one for clients)
an external thing that should have a local name ExternalName
expose to the outside, cloud LoadBalancer (or one Ingress/Gateway for many)
expose to the outside, bare metal LoadBalancer + MetalLB, or NodePort
the backend needs the real client IP externalTrafficPolicy: Local
(Ingress and Gateway, 16.19 and 16.26, put many HTTP sites behind one address.)
What you can now do:
- expose a Service outside the cluster with NodePort or LoadBalancer
- keep the client's real IP with externalTrafficPolicy: Local, and know its cost
- name an outside system with ExternalName, and find individual pods with headless