$ kubectl logs frontend -n checkout --tail=3
checkout call failed
wget: can't connect to remote host (10.96.142.147): Connection refused
checkout call failedThe checkout pods are Running and Ready, with zero restarts. Their logs are clean. Someone has redeployed twice. Every call through the Service is still refused, instantly.
Redeploying changes nothing, because the pods are fine. The problem is the wiring between the Service and the pods.
What is actually happening
A Service is a stable virtual IP (the ClusterIP) plus a selector, a set of labels. The endpoints controller keeps an EndpointSlice listing every pod that matches the selector, with each pod's IP, the target port and whether the pod is ready. kube-proxy on every node turns that list into rules that rewrite connections to the ClusterIP into connections to one ready pod IP and port.
Three links in that chain can break, and each one has its own fingerprint:
| Broken link | EndpointSlice | What callers get |
|---|---|---|
| the selector matches no pod | empty (<unset>) | Connection refused, instantly: kube-proxy REJECTs a Service with no endpoints |
| pods selected, but none Ready | the pods, ready: false | the same refusal: not-ready pods get no traffic |
endpoints fine, wrong targetPort | pod IPs on the wrong port | Connection refused from the pod itself: nothing listens there |
"Refused" means you reached something that said no. A timeout would point somewhere else: a NetworkPolicy, a port the Service doesn't expose, an unreachable node. Put a proxy in front and the same empty Service turns into a 502 or 503 instead: an nginx 502 Bad Gateway, or on OpenShift the router's "Application is not available" page.
Diagnosis
1. Look at the endpoints first
$ kubectl describe svc checkout -n checkout
Name: checkout
Namespace: checkout
Labels: app.kubernetes.io/name=checkout
Selector: app=checkout
Type: ClusterIP
IP: 10.96.142.147
Port: http 80/TCP
TargetPort: http/TCP
Endpoints:
...
$ kubectl get endpointslice -n checkout -l kubernetes.io/service-name=checkout
NAME ADDRESSTYPE PORTS ENDPOINTS AGE
checkout-cxd2k IPv4 80 <unset> 7m1sEndpoints: is empty. The Service selects nothing.
2. Empty: put the selector next to the labels
$ kubectl get svc checkout -n checkout -o jsonpath='{.spec.selector}'; echo
{"app":"checkout"}
$ kubectl get pods -n checkout --show-labels
NAME READY STATUS RESTARTS AGE LABELS
checkout-x9xsxln6wg-8gfwg 1/1 Running 0 7m1s app.kubernetes.io/name=checkout,app.kubernetes.io/part-of=shop,app.kubernetes.io/version=3.2.0,pod-template-hash=x9xsxln6wgThe deploy moved the pods to the recommended app.kubernetes.io/* labels. The Service still asks for app=checkout. Replace the whole selector with a JSON patch:
$ kubectl patch svc checkout -n checkout --type=json -p='[{"op":"replace","path":"/spec/selector","value":{"app.kubernetes.io/name":"checkout"}}]'
service/checkout patched
$ kubectl get endpointslice -n checkout -l kubernetes.io/service-name=checkout
NAME ADDRESSTYPE PORTS ENDPOINTS AGE
checkout-cxd2k IPv4 8080 10.244.2.127,10.244.1.16 7m5sA plain kubectl patch -p (a merge patch) would merge the maps. You'd get both keys in the selector, matching nothing.
3. Pods listed but not ready: it's the readiness probe
Another team, same symptom, but this time they added a readiness probe in the morning's release:
$ kubectl get pods -n search
NAME READY STATUS RESTARTS AGE
search-rwbp59mddf-p85mn 0/1 Running 0 4m2s
search-rwbp59mddf-tfbk5 0/1 Running 0 4m2s
search-rwbp59mddf-vx7mt 0/1 Running 0 4m2s
$ kubectl get endpointslice -n search -o yaml | grep -A3 conditions
conditions:
ready: false
serving: false
terminating: false0/1 Running is the tell. A readiness probe is a check the kubelet runs against the app (here an HTTP GET every 5 seconds). While it fails, the pod stays in the EndpointSlice with ready: false and gets no traffic. It is never restarted: readiness only switches traffic. (A failing liveness probe restarts the container, and that ends in CrashLoopBackOff.) The reason is in the events:
$ kubectl describe pod -n search -l app=search | grep -m1 "Readiness probe failed"
Warning Unhealthy 4s (x48 over 3m59s) kubelet Readiness probe failed: HTTP probe failed with statuscode: 404The probe asks for /readyz and the app answers 404: there's no such path. This app's readiness endpoint is /ready. Fix the path in the pod template. A rollout follows, and three ready endpoints appear.
4. Endpoints there, still refused: compare the three ports
$ kubectl get endpoints ledger-api -n ledger
NAME ENDPOINTS AGE
ledger-api 10.244.2.201:80,10.244.1.53:80 5m2s
$ kubectl get deploy ledger-api -n ledger -o jsonpath='{.spec.template.spec.containers[0].ports}'; echo
[{"name":"http","containerPort":8080,"protocol":"TCP"}]The Service sends to port 80, but the new image listens on 8080. Endpoints only carry the Service's idea of the port. They're no proof that anything listens there. Prove it from a pod: wget <pod-ip>:8080 answers and :80 is refused. Then point targetPort at the port's name:
$ kubectl patch svc ledger-api -n ledger --type=json -p='[{"op":"replace","path":"/spec/ports/0/targetPort","value":"http"}]'
service/ledger-api patchedNow the pod spec says "my http port is 8080" and the Service says "send to http". The next port change only has to happen in one place.
The order of checks
kubectl describe svc: isEndpointsempty? Then compare the selector with--show-labels.- Are the endpoints there but pods
0/1? Then read the readiness probe events. - Are the endpoints there and pods Ready? Then compare the endpoint port with
containerPortand test the pod IP directly.
How to prevent it
- Change labels on the Deployment, the Service, NetworkPolicies and anything else that selects pods in one commit.
- Use named target ports (
targetPort: http). - Before you add a readiness probe, test its path against a running pod. Readiness also decides when a rolling update moves on, so a correct probe matters twice (downtime on every deploy).
- Smoke-test through the Service after a deploy, not against a pod. "Pods are Running" is not "pods get traffic".
Practise it
Chapter 16 has all three as incidents: "checkout refuses every connection", "the endpoints are there, and it still refuses" and "Running, zero restarts, and no traffic". Each one ends with you patching the right object, with no redeploy.