The problem: pods keep changing address
Your shop has three copies of a web app running as pods. Another app (say, the checkout) needs to call it. Which IP does it call? A pod's IP changes every time the pod is replaced - after a crash, a rollout, a node going away. A hard-coded pod IP is broken within the hour.
In chapter 9 you met the fix for "many servers, one address": a load balancer (9.23) - clients talk to one address and it forwards each connection to one of the servers behind it. Kubernetes builds one in: the Service. This chapter is about how clients find and reach pods (Services, DNS, Ingress, network rules) and where pods keep data that must outlive them (volumes). This first lesson is the Service itself.
What you need to know already: pods, Deployments and ReplicaSets (15.14, 15.16), labels and selectors (15.26), namespaces (15.26), kubectl get / -o wide / jsonpath (15.38), ports (9.8), load balancers (9.23), DNS names (8.16).
Setting up: a few pods to put behind a Service
The lessons from here on use the speed kit from 15.3: k is short for kubectl. If this shell has no k (you skipped 15.4), define it now:
$ alias k=kubectl
The lesson's shop namespace, built in four commands. The app is agnhost - a small test web server from the Kubernetes project; with netexec it serves HTTP on port 8080.
$ k delete ns shop --ignore-not-found
$ k create ns shop
$ k create deployment web -n shop --image=registry.k8s.io/e2e-test-images/agnhost:2.53 --replicas=3 --port=8080 -- /agnhost netexec --http-port=8080
$ k patch deploy web -n shop --type=json -p '[{"op":"add","path":"/spec/template/spec/containers/0/ports/0/name","value":"http"}]'
What each one does:
k delete ns shop --ignore-not-found remove chapter 15's shop namespace (and everything in it);
--ignore-not-found = no error if it is not there.
kubectl waits until it is fully gone.
k create ns shop a fresh, empty namespace
k create deployment web ... a Deployment "web": --image which image, --replicas=3
three pods, --port=8080 declares the container port;
everything after -- is the container's command
k patch deploy web --type=json -p change one field in place (15.40): give port 8080
the NAME "http" (you will see why below)
Now watch a pod's IP change. -o wide adds the IP and NODE columns:
$ k get pods -n shop -o wide
NAME READY STATUS RESTARTS AGE IP NODE
web-b9tlh6zz5x-77n2f 1/1 Running 0 60s 10.244.2.164 worker-2
web-b9tlh6zz5x-7vd6w 1/1 Running 0 60s 10.244.1.16 worker-1
web-b9tlh6zz5x-mg9nm 1/1 Running 0 60s 10.244.2.201 worker-2
$ k delete pod -n shop $(k get pods -n shop -o wide | awk '/worker-1/{print $1; exit}')
pod "web-b9tlh6zz5x-7vd6w" deleted from shop namespace
$ k get pods -n shop -o wide
NAME READY STATUS RESTARTS AGE IP NODE
web-b9tlh6zz5x-77n2f 1/1 Running 0 65s 10.244.2.164 worker-2
web-b9tlh6zz5x-mg9nm 1/1 Running 0 65s 10.244.2.201 worker-2
web-b9tlh6zz5x-q2w8x 1/1 Running 0 4s 10.244.1.93 worker-1
The $(...) picks the name of the pod on worker-1 (awk: the first line mentioning worker-1, print column 1, stop). You deleted that pod; the ReplicaSet made a replacement (q2w8x, AGE 4s) and it has a new IP, 10.244.1.93 instead of 10.244.1.16. Every pod gets its IP from its node's slice of the pod address range (16.35 shows where those ranges come from).
Nothing that talks to web can hard-code 10.244.1.16.
The Service object
A Service is the fix: one IP and one DNS name that stay put for as long as the Service exists, plus a label selector (15.26) that says which pods are behind it right now. Clients use the Service; the Service keeps track of the pods.
The Service's IP is called its ClusterIP: a virtual IP (no machine actually owns it - 16.3 shows how it works anyway) that is reachable only from inside the cluster. ClusterIP is also the name of the default Service type.
apiVersion: v1
kind: Service
metadata:
name: web
namespace: shop
spec:
selector:
app: web # every Ready pod with this label, in THIS namespace
ports:
- name: http # optional with one port, required with several
port: 80 # what clients connect to: web:80
targetPort: http # the pod's port - a number, or a named containerPort
protocol: TCP # the default
Read it as: "a Service called web in namespace shop; its pods are the ones labelled app=web; clients connect to port 80 and the traffic goes to the pods' port called http".
Three port numbers people mix up
port the Service's port. Clients use web:80.
targetPort where the traffic lands in the pod. Defaults to port.
containerPort the port listed in the pod spec (ports: - containerPort: 8080).
It is documentation, plus a NAME a Service can target.
Nothing enforces it: an app listening on 9000 with containerPort
8080 simply gets connections to 8080 refused.
Compare it with Docker's -p 8080:80 (11.15): the outside number and the inside number are different things. Here port is the outside number, targetPort the inside one.
A named targetPort (targetPort: http) is looked up in each pod: it means "whatever port this pod calls http". During a rollout where v1 listens on 8080 and v2 on 9090, both named http, the Service reaches each on its own port. With a number, one of them breaks.
Creating one
kubectl expose builds a Service from an existing object. svc is the short name for services in any kubectl command.
$ k expose deploy web -n shop --port=80 --target-port=http
service/web exposed
$ k get svc -n shop
NAME TYPE CLUSTER-IP EXTERNAL-IP PORT(S) AGE
web ClusterIP 10.96.142.147 <none> 80/TCP 5s
| part | means |
|---|---|
expose deploy web | make a Service for Deployment web (same name, same namespace) |
--port=80 | the Service's port |
--target-port=http | the pods' port, here by name |
The columns of get svc:
NAME the Service's name - also its DNS name inside the cluster
TYPE ClusterIP (inside only); 16.6 covers the other types
CLUSTER-IP the stable virtual IP
EXTERNAL-IP an address reachable from outside the cluster; <none> for ClusterIP
PORT(S) port/protocol clients use (80/TCP)
expose copies the Deployment's selector (app=web) into the Service - not its labels. The ClusterIP comes from the service CIDR: an address range (8.3) reserved for Services, set when the cluster was built. The API server got it as --service-cluster-ip-range=10.96.0.0/12 on this cluster (kubeadm, the tool that built it, uses that range by default).
There are also generators that write a Service from scratch:
k create service clusterip web --tcp=80:8080 -n shop # selector app=web (!)
k create service nodeport web --tcp=80:8080 --node-port=30900
k create service clusterip db --clusterip=None --tcp=5432:5432 # headless (15.19)
--tcp=80:8080 means port 80, targetPort 8080. create service always sets selector: app=<name> - fine if your pods carry that label, silently empty if they do not.
EndpointSlices: who is behind it right now
An endpoint is one address traffic can go to: a pod IP plus a port. A controller in the control plane (15.5) watches Services and Pods and writes the list of matching pod IPs into EndpointSlice objects. It also still writes the older, one-object-per-Service Endpoints object, kept for compatibility. (A "slice" because a Service with thousands of pods gets several of these objects, up to 100 endpoints each.)
$ k get endpointslice -n shop -l kubernetes.io/service-name=web
NAME ADDRESSTYPE PORTS ENDPOINTS AGE
web-5mp8g IPv4 8080 10.244.2.164,10.244.1.16,10.244.2.201 60s
-l kubernetes.io/service-name=web selects by label (15.26): every slice carries the name of its Service in that label, and the slice's own name gets a random suffix, so the label is how you find them. The columns:
ADDRESSTYPE IPv4 or IPv6
PORTS the resolved targetPort: the name http became the number 8080
ENDPOINTS the pod IPs behind the Service right now
The full object (-o yaml, 15.38) shows what kube-proxy on each node uses to route traffic. The inner $(... -o name) prints the slice as endpointslice.discovery.k8s.io/web-5mp8g, so you do not have to type the random suffix:
$ k get $(k get endpointslice -n shop -l kubernetes.io/service-name=web -o name) -n shop -o yaml
addressType: IPv4
apiVersion: discovery.k8s.io/v1
endpoints:
- addresses:
- 10.244.2.164
conditions:
ready: true # counts for traffic
serving: true # would serve (true even while terminating)
terminating: false
nodeName: worker-2
targetRef:
kind: Pod
name: web-b9tlh6zz5x-77n2f
...
ports:
- port: 8080
protocol: TCP
Each entry under endpoints: is one pod: its address, the node it runs on (nodeName), which pod it is (targetRef), and three conditions.
Only ready: true endpoints receive traffic. A pod is Ready when its containers are up and pass their health check (a readiness probe: a check the kubelet runs against the app, 15.14 showed the READY column). A pod that is Running but not Ready - its check fails, or it is shutting down - stays in the slice with ready: false, and gets nothing. Readiness is the switch that takes a pod out of every Service that selects it.
Later (Ch 17): writing readiness and liveness probes yourself.
The older Endpoints object shows the same list in one line, IP:port:
$ k get endpoints web -n shop
NAME ENDPOINTS AGE
web 10.244.2.164:8080,10.244.1.16:8080,10.244.2.201:8080 60s
describe svc puts it all together. The lines that matter: Selector (which pods), IP (the ClusterIP), Port and TargetPort, and Endpoints (who is Ready behind it now):
$ k describe svc web -n shop
Name: web
Namespace: shop
Labels: app=web
Annotations: <none>
Selector: app=web
Type: ClusterIP
IP Family Policy: SingleStack
IP Families: IPv4
IP: 10.96.142.147
IPs: 10.96.142.147
Port: <unset> 80/TCP
TargetPort: http/TCP
Endpoints: 10.244.1.16:8080,10.244.2.164:8080,10.244.2.201:8080
Session Affinity: None
Internal Traffic Policy: Cluster
Events: <none>
Session Affinity and Internal Traffic Policy are routing options covered in 16.6; IP Family lines say IPv4 only. An empty Endpoints: line is the first thing to look for when "the service does not work".
Using it from a pod
A Service is reached from inside the cluster, so you need a client pod. Two of them, one in shop and one in another namespace. k run starts a single pod; busybox is a tiny image with wget, nc, nslookup and a shell, and sleep 1d keeps it alive for a day so you can exec into it. Create the namespace a few seconds before run: a brand-new namespace has no default ServiceAccount (the identity every pod runs as) for a moment, and the pod is refused.
$ k delete ns other --ignore-not-found; k create ns other
$ k run toolbox -n shop --image=busybox:1.36 -- sleep 1d
$ k run toolbox -n other --image=busybox:1.36 -- sleep 1d
$ k exec -n shop toolbox -- wget -qO- web/hostname
web-b9tlh6zz5x-mg9nm
$ k exec -n shop toolbox -- wget -qO- web.shop.svc.cluster.local/hostname
web-b9tlh6zz5x-77n2f
$ k exec -n other toolbox -- wget -qO- web.shop/hostname # another namespace
web-b9tlh6zz5x-7vd6w
k exec -n shop toolbox -- CMD runs CMD inside the toolbox pod (15.15). wget -qO- URL fetches a URL: -q quiet (no progress lines), -O- write the page to the screen instead of a file.
Three names reached the same Service:
web from the same namespace
web.shop <service>.<namespace>, from any namespace
web.shop.svc.cluster.local the full name (16.13 explains every part)
That is DNS (8.16): the cluster's DNS server, CoreDNS (15.7), answers these names with the ClusterIP.
The agnhost app (registry.k8s.io/e2e-test-images/agnhost:2.53, the image Kubernetes' own tests use) answers /hostname with the pod's name and /clientip with the address the connection came from. Ask a few times and different pods answer: kube-proxy picks an endpoint per connection, not per request.
Pods started after a Service also get environment variables for it. This is an old convention copied from Docker's container links, still on by default. env | grep WEB_ lists the environment and keeps the lines with WEB_:
$ k exec -n shop toolbox -- env | grep WEB_
WEB_SERVICE_HOST=10.96.142.147
WEB_SERVICE_PORT=80
WEB_PORT=tcp://10.96.142.147:80
WEB_PORT_80_TCP=tcp://10.96.142.147:80
...
Only for Services that existed when the pod started, and only in its own namespace. So use DNS names, not these variables. In namespaces with hundreds of Services, set enableServiceLinks: false in the pod spec: the variable block gets huge and slows container start.
The failure you will see most
$ k exec -n shop toolbox -- wget -qO- -T 2 web:8080/
wget: download timed out
command terminated with exit code 1
$ k exec -n shop toolbox -- nc -zv -w 2 web 80
web (10.96.142.147:80) open
wget -T 2 gives up after 2 seconds. nc -zv -w 2 web 80 tests a TCP port without sending data (9.1): -z just connect, -v say what happened, -w 2 wait at most 2 seconds.
web:8080 times out: the Service has no port 8080, so no rule matches and the packet wanders off to the default gateway (8.11). Connect to the Service port, not the container port. A Service with no endpoints fails differently - an immediate Connection refused - which is the next lesson's subject.
What you can now do:
- put a Service in front of pods with
kubectl expose, and pick port and targetPort - check who is behind it with
get endpointsliceanddescribe svc - reach it from a pod by name, from the same or another namespace