Node NotReady
The problem. The second half of the catalogue: failures that are not about one pod but about a node, a Service or DNS - where the pod looks fine and still nothing works.
What you need to know already: the pod catalogue (18.28), node lifecycle (18.25), the kubelet journal (18.1), Services and endpoints (16.1), CoreDNS and resolv.conf (16.13), NetworkPolicy (16.29), readiness probes (17.20).
Symptom: kubectl get nodes NotReady; pods on it stop being ready, new pods avoid it (taints), after 5 minutes its pods are evicted (18.25).
Diagnosis: the Ready condition's Status and message tell you which kind:
$ kubectl describe node worker-1 | grep -E '^\s+Ready'
Ready False ... KubeletNotReady container runtime is down, PLEG is not healthy: pleg was last seen active 3m21.134s ago; threshold is 3m0s
| Ready | message says | on the node | fix |
|---|---|---|---|
| Unknown | Kubelet stopped posting node status | systemctl status kubelet: inactive / activating (auto-restart) | start it; if it crash-loops, journalctl -u kubelet names the file or flag |
| False | container runtime is down, PLEG is not healthy | systemctl status containerd: inactive | systemctl start containerd, then find why it died (journalctl -u containerd) |
| False | container runtime network not ready ... cni plugin not initialized | ls /etc/cni/net.d empty | restart calico-node on that node (it rewrites the config), or re-apply the CNI |
| True + DiskPressure True | kubelet has disk pressure | df -h / > 85-90% | free space: logs, dumps, crictl rmi --prune, journalctl --vacuum-size |
| Unknown, kubelet running | (journal) dial tcp 10.64.0.10:6443: connect: connection refused or x509 | the node is fine | the control plane is broken - go to cp-1 |
Two things people get wrong about disk: DiskPressure does not make the node NotReady - it stays Ready with the node.kubernetes.io/disk-pressure:NoSchedule taint, the kubelet garbage-collects images and then evicts pods, BestEffort first:
# an illustration: the incidents in this chapter build each of these
kubectl get pods -n shop
NAME READY STATUS RESTARTS AGE
reports-5b8-q2x7c 0/1 Evicted 0 14m
kubectl describe pod reports-5b8-q2x7c | grep -A2 Message
Message: The node was low on resource: ephemeral-storage. Threshold quantity: 1949574Ki, available: 584920Ki.
and evicted pods are not cleaned up - they stay as Evicted (phase Failed) until the pod garbage collector reaches its threshold or you delete them. Find where the space went on the node, the Ch 4 way (4.19):
learner@worker-2:~$ df -h /
Filesystem Size Used Avail Use% Mounted on
/dev/root 19G 18G 1.1G 95% /
learner@worker-2:~$ sudo du -xh / --max-depth=3 2>/dev/null | sort -h | tail -5
The usual suspects on a node: container images (/var/lib/containerd, sudo crictl rmi --prune removes images no container uses), the journal (/var/log/journal, sudo journalctl --vacuum-size=500M), and something someone left: heap dumps, core files, a debug log in /var/log. Container logs in /var/log/pods are rotated by the kubelet (containerLogMaxSize 10Mi by default) and are rarely it.
Service has no endpoints
Symptom: curl svc refuses or times out, the Service exists, pods look fine.
Diagnosis: does the Service have endpoints, and why not:
# an illustration: the incidents in this chapter build each of these
kubectl get endpointslices -l kubernetes.io/service-name=front
NAME ADDRESSTYPE PORTS ENDPOINTS AGE
front-bwmtp IPv4 8080 <unset> 19s
kubectl describe svc front | grep -E 'Selector|Endpoints'
Selector: app=frontend
Endpoints:
kubectl get pods --show-labels
NAME READY STATUS RESTARTS AGE LABELS
front-5d9c7b6f4-8xkq2 1/1 Running 0 2m app=front,pod-template-hash=5d9c7b6f4
Two causes, one check each:
- The selector matches no pods -
app=frontendvsapp=front. Comparedescribe svcSelector withget pods --show-labels, or test the selector directly:kubectl get pods -l app=frontendreturns nothing. Fix the Service's selector (or the pod labels - but the Deployment's selector is immutable). - The pods match but are not Ready - a failing readiness probe keeps them out of the ready endpoints.
kubectl get podsshows0/1 Running; describe showsReadiness probe failed: HTTP probe failed with statuscode: 404. Fix the probe (or the app).
And a third that looks like the same thing: endpoints exist but targetPort is wrong - connections are refused on the pod IP. describe svc TargetPort vs the container's port.
DNS failing inside a pod
Symptom: the app logs UnknownHostException / no such host; from a pod:
# an illustration: the incidents in this chapter build each of these
kubectl exec t -- nslookup kubernetes.default
;; connection timed out; no servers could be reached
Diagnosis, from the outside in:
# an illustration: the incidents in this chapter build each of these
kubectl exec t -- cat /etc/resolv.conf # which server does the pod ask?
search default.svc.cluster.local svc.cluster.local cluster.local
nameserver 10.96.0.10
options ndots:5
$ kubectl get svc -n kube-system kube-dns # is that the DNS Service IP?
$ kubectl get endpointslices -n kube-system -l kubernetes.io/service-name=kube-dns # any CoreDNS behind it?
$ kubectl get pods -n kube-system -l k8s-app=kube-dns # are they running?
$ kubectl logs -n kube-system -l k8s-app=kube-dns # what CoreDNS says
| finding | cause |
|---|---|
| timed out, no kube-dns endpoints | CoreDNS scaled to 0 / crash-looping / Pending |
| timed out, CoreDNS fine | a NetworkPolicy blocks egress to port 53 (chapter 16) |
| nameserver is not the kube-dns ClusterIP | the node's kubelet has a wrong clusterDNS in /var/lib/kubelet/config.yaml (only pods created after the change get it) |
| NXDOMAIN for an external name, internal fine | CoreDNS's forward upstream (or the node's resolv.conf) |
| slow but works | ndots:5 search expansion (chapters 8 and 16): use FQDNs with a trailing dot |
timed out and NXDOMAIN are different failures: the first means nobody answered, the second means somebody answered "no such name".
Expired certificates
Covered in 18.11. The one-line reminder: x509: certificate has expired from kubectl = the apiserver's serving cert; Unauthorized = your client cert. sudo kubeadm certs check-expiration on the control plane, renew, restart the static pods, re-copy admin.conf.
The whole catalogue on one screen
| symptom | likely causes | first commands |
|---|---|---|
| CrashLoopBackOff | bad config, missing env, dependency down, wrong command, exit 0 in a Deployment | describe Last State, logs --previous |
| ImagePullBackOff / ErrImagePull | wrong tag, no imagePullSecret, registry unreachable, rate limit | describe Events: the registry's answer |
| Pending | requests do not fit, no matching node, untolerated taint, unbound PVC, no scheduler | describe FailedScheduling (or no events), get pvc, get sc |
| ContainerCreating (stuck) | FailedMount, CNI missing, runtime down | describe Events, the node |
| Terminating forever | finalizer, node gone | -o jsonpath={.metadata.finalizers}, the node |
| OOMKilled / 137 | limit too low, JVM heap vs limit, leak | describe Last State, top pod |
| Node NotReady | kubelet stopped, runtime down, CNI missing, control plane unreachable | describe node Conditions, systemctl status kubelet, journalctl -u kubelet |
| Node DiskPressure / Evicted pods | full root disk: images, journal, dumps | df -h, du, crictl rmi --prune |
| Service no endpoints | selector != labels, pods not Ready | get endpointslices, describe svc, get pods --show-labels |
| DNS failing in a pod | CoreDNS down, NetworkPolicy on 53, wrong clusterDNS | nslookup from a pod, CoreDNS pods and logs |
| kubectl: connection refused | apiserver down (manifest, etcd, certs) | ssh cp-1, crictl ps -a, crictl logs, journalctl -u kubelet |
| kubectl: x509 expired | certificates | kubeadm certs check-expiration |
| everything broke after an upgrade | removed API version, skew violated, not restarted | kubectl api-resources, get nodes VERSION |