OnCallReady

Lesson 18.34 · Kubernetes: Cluster Operations & Troubleshooting · 15 min read

The failure catalogue II: nodes, Services, DNS, certificates

In plain words

Think of a plumber's checklist for "no water in the flat". Is the building's main valve open? Is the pipe to this flat blocked? Is the tap itself broken? Is the water meter reading the wrong flat? Each check has its own tool, and a good plumber goes through them in order instead of replacing random pipes.

This lesson is that checklist for the layers under pods. Node NotReady: the Ready condition's message says whether it's the kubelet, containerd, the CNI or the control plane. DiskPressure: the node stays Ready but evicts pods. Service without endpoints: selector versus labels, or pods not Ready. DNS failing: timed out means nobody answered, NXDOMAIN means somebody said no. Certificates: x509 versus Unauthorized.

Node NotReady

The problem. The second half of the catalogue: failures that are not about one pod but about a node, a Service or DNS - where the pod looks fine and still nothing works.

What you need to know already: the pod catalogue (18.28), node lifecycle (18.25), the kubelet journal (18.1), Services and endpoints (16.1), CoreDNS and resolv.conf (16.13), NetworkPolicy (16.29), readiness probes (17.20).

Symptom: kubectl get nodes NotReady; pods on it stop being ready, new pods avoid it (taints), after 5 minutes its pods are evicted (18.25).

Diagnosis: the Ready condition's Status and message tell you which kind:

$ kubectl describe node worker-1 | grep -E '^\s+Ready'
  Ready            False    ...   KubeletNotReady     container runtime is down, PLEG is not healthy: pleg was last seen active 3m21.134s ago; threshold is 3m0s
Readymessage sayson the nodefix
UnknownKubelet stopped posting node statussystemctl status kubelet: inactive / activating (auto-restart)start it; if it crash-loops, journalctl -u kubelet names the file or flag
Falsecontainer runtime is down, PLEG is not healthysystemctl status containerd: inactivesystemctl start containerd, then find why it died (journalctl -u containerd)
Falsecontainer runtime network not ready ... cni plugin not initializedls /etc/cni/net.d emptyrestart calico-node on that node (it rewrites the config), or re-apply the CNI
True + DiskPressure Truekubelet has disk pressuredf -h / > 85-90%free space: logs, dumps, crictl rmi --prune, journalctl --vacuum-size
Unknown, kubelet running(journal) dial tcp 10.64.0.10:6443: connect: connection refused or x509the node is finethe control plane is broken - go to cp-1

Two things people get wrong about disk: DiskPressure does not make the node NotReady - it stays Ready with the node.kubernetes.io/disk-pressure:NoSchedule taint, the kubelet garbage-collects images and then evicts pods, BestEffort first:

# an illustration: the incidents in this chapter build each of these
kubectl get pods -n shop
NAME                  READY   STATUS    RESTARTS   AGE
reports-5b8-q2x7c     0/1     Evicted   0          14m
kubectl describe pod reports-5b8-q2x7c | grep -A2 Message
Message:          The node was low on resource: ephemeral-storage. Threshold quantity: 1949574Ki, available: 584920Ki.

and evicted pods are not cleaned up - they stay as Evicted (phase Failed) until the pod garbage collector reaches its threshold or you delete them. Find where the space went on the node, the Ch 4 way (4.19):

learner@worker-2:~$ df -h /
Filesystem      Size  Used Avail Use% Mounted on
/dev/root        19G   18G  1.1G  95% /
learner@worker-2:~$ sudo du -xh / --max-depth=3 2>/dev/null | sort -h | tail -5

The usual suspects on a node: container images (/var/lib/containerd, sudo crictl rmi --prune removes images no container uses), the journal (/var/log/journal, sudo journalctl --vacuum-size=500M), and something someone left: heap dumps, core files, a debug log in /var/log. Container logs in /var/log/pods are rotated by the kubelet (containerLogMaxSize 10Mi by default) and are rarely it.

Service has no endpoints

Symptom: curl svc refuses or times out, the Service exists, pods look fine.

Diagnosis: does the Service have endpoints, and why not:

# an illustration: the incidents in this chapter build each of these
kubectl get endpointslices -l kubernetes.io/service-name=front
NAME          ADDRESSTYPE   PORTS   ENDPOINTS   AGE
front-bwmtp   IPv4          8080    <unset>     19s
kubectl describe svc front | grep -E 'Selector|Endpoints'
Selector:                 app=frontend
Endpoints:
kubectl get pods --show-labels
NAME                     READY   STATUS    RESTARTS   AGE   LABELS
front-5d9c7b6f4-8xkq2    1/1     Running   0          2m    app=front,pod-template-hash=5d9c7b6f4

Two causes, one check each:

  1. The selector matches no pods - app=frontend vs app=front. Compare describe svc Selector with get pods --show-labels, or test the selector directly: kubectl get pods -l app=frontend returns nothing. Fix the Service's selector (or the pod labels - but the Deployment's selector is immutable).
  2. The pods match but are not Ready - a failing readiness probe keeps them out of the ready endpoints. kubectl get pods shows 0/1 Running; describe shows Readiness probe failed: HTTP probe failed with statuscode: 404. Fix the probe (or the app).

And a third that looks like the same thing: endpoints exist but targetPort is wrong - connections are refused on the pod IP. describe svc TargetPort vs the container's port.

DNS failing inside a pod

Symptom: the app logs UnknownHostException / no such host; from a pod:

# an illustration: the incidents in this chapter build each of these
kubectl exec t -- nslookup kubernetes.default
;; connection timed out; no servers could be reached

Diagnosis, from the outside in:

# an illustration: the incidents in this chapter build each of these
kubectl exec t -- cat /etc/resolv.conf                  # which server does the pod ask?
search default.svc.cluster.local svc.cluster.local cluster.local
nameserver 10.96.0.10
options ndots:5
$ kubectl get svc -n kube-system kube-dns                 # is that the DNS Service IP?
$ kubectl get endpointslices -n kube-system -l kubernetes.io/service-name=kube-dns   # any CoreDNS behind it?
$ kubectl get pods -n kube-system -l k8s-app=kube-dns     # are they running?
$ kubectl logs -n kube-system -l k8s-app=kube-dns         # what CoreDNS says
findingcause
timed out, no kube-dns endpointsCoreDNS scaled to 0 / crash-looping / Pending
timed out, CoreDNS finea NetworkPolicy blocks egress to port 53 (chapter 16)
nameserver is not the kube-dns ClusterIPthe node's kubelet has a wrong clusterDNS in /var/lib/kubelet/config.yaml (only pods created after the change get it)
NXDOMAIN for an external name, internal fineCoreDNS's forward upstream (or the node's resolv.conf)
slow but worksndots:5 search expansion (chapters 8 and 16): use FQDNs with a trailing dot

timed out and NXDOMAIN are different failures: the first means nobody answered, the second means somebody answered "no such name".

Expired certificates

Covered in 18.11. The one-line reminder: x509: certificate has expired from kubectl = the apiserver's serving cert; Unauthorized = your client cert. sudo kubeadm certs check-expiration on the control plane, renew, restart the static pods, re-copy admin.conf.

The whole catalogue on one screen

symptomlikely causesfirst commands
CrashLoopBackOffbad config, missing env, dependency down, wrong command, exit 0 in a Deploymentdescribe Last State, logs --previous
ImagePullBackOff / ErrImagePullwrong tag, no imagePullSecret, registry unreachable, rate limitdescribe Events: the registry's answer
Pendingrequests do not fit, no matching node, untolerated taint, unbound PVC, no schedulerdescribe FailedScheduling (or no events), get pvc, get sc
ContainerCreating (stuck)FailedMount, CNI missing, runtime downdescribe Events, the node
Terminating foreverfinalizer, node gone-o jsonpath={.metadata.finalizers}, the node
OOMKilled / 137limit too low, JVM heap vs limit, leakdescribe Last State, top pod
Node NotReadykubelet stopped, runtime down, CNI missing, control plane unreachabledescribe node Conditions, systemctl status kubelet, journalctl -u kubelet
Node DiskPressure / Evicted podsfull root disk: images, journal, dumpsdf -h, du, crictl rmi --prune
Service no endpointsselector != labels, pods not Readyget endpointslices, describe svc, get pods --show-labels
DNS failing in a podCoreDNS down, NetworkPolicy on 53, wrong clusterDNSnslookup from a pod, CoreDNS pods and logs
kubectl: connection refusedapiserver down (manifest, etcd, certs)ssh cp-1, crictl ps -a, crictl logs, journalctl -u kubelet
kubectl: x509 expiredcertificateskubeadm certs check-expiration
everything broke after an upgraderemoved API version, skew violated, not restartedkubectl api-resources, get nodes VERSION

Why it helps

These are the incidents that affect many workloads at once, so they're the ones that page you. A node NotReady with "PLEG is not healthy" is containerd; with "cni plugin not initialized" it's the CNI agent; with a kubelet that's running but logging connection refused to 6443, it's the control plane and every node will be NotReady soon. Knowing the mapping saves you from fixing the wrong machine.

DiskPressure is subtle and common: the node looks Ready, pods get evicted and pile up as Evicted, and deleting them does nothing while the disk is still full of images and journals. And "the service doesn't work" or "DNS is broken" tickets come in weekly; the one-screen catalogue at the end of this lesson is what you'll carry into interviews and the admin exam.

Commands in this lesson

kubectl

FAQ

Does DiskPressure make the node NotReady?

No. The node stays Ready with the node.kubernetes.io/disk-pressure:NoSchedule taint. The kubelet garbage-collects unused images and then evicts pods, BestEffort first, and the evicted pods stay listed as Evicted until cleaned up. Free the space on the node: unused images (crictl rmi --prune), the journal (journalctl --vacuum-size), and leftover dumps or debug logs.

A node is NotReady but the kubelet is running. What's wrong?

Read the kubelet journal. If it logs dial tcp <cp-ip>:6443: connect: connection refused or x509 errors when updating node status, the node is fine and the control plane isn't: the API server is down or its certificates expired. Go to the control-plane node. If it logs runtime or CNI errors, the problem is on this node.

What's the difference between "timed out" and NXDOMAIN in DNS?

"connection timed out; no servers could be reached" means nobody answered: CoreDNS is down or scaled to zero, or a NetworkPolicy drops UDP 53. NXDOMAIN means a DNS server answered "no such name": DNS works, and the name, the search path or the upstream forwarder is wrong. They send you to completely different places.

Can I fix a Service selector mismatch by relabelling the pods?

For pods owned by a Deployment, not easily: the Deployment's selector is immutable, and changing the template labels to something the selector doesn't match is rejected. It's usually simpler to fix the Service's selector to match the pods' labels. Test it first with kubectl get pods -l <selector>.

Why do pods created before a kubelet fix still have the wrong nameserver?

The kubelet writes a pod's /etc/resolv.conf when the pod is created, from its clusterDNS setting in /var/lib/kubelet/config.yaml. Fixing the kubelet config (and restarting the kubelet) only affects pods created afterwards; existing pods keep the old file until they're recreated.

In an interview Junior

A Service has no endpoints. How do you find out why?

The pods look fine and curl to the Service fails. Check what is behind it:

kubectl get endpointslices -l kubernetes.io/service-name=front
kubectl describe svc front
kubectl get pods --show-labels

Two causes, one check each:

  1. The selector matches no pods - app=frontend in the Service, app=front on the pods. Test it: kubectl get pods -l app=frontend returns nothing. Fix the Service's selector (a Deployment's selector is immutable).
  2. The pods match but are not Ready - 0/1 Running, and describe pod shows Readiness probe failed. Only Ready pods become ready endpoints. Fix the probe or the app.

And the look-alike: endpoints exist but the targetPort is wrong, so connections to the pod IP are refused - compare TargetPort with the container's port.

Also asked: Pods are being evicted with "The node was low on resource: ephemeral-storage". What do you do? · DNS fails inside every pod. How do you debug it? · What is the difference between a DNS timeout and NXDOMAIN?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.