In plain words
Imagine the lights go out in your house. You don't start by unscrewing every bulb. First: is it just your house or the whole street? If the whole street, it's the power company. If just your house, check the fuse box. If only one room, check that room's switch and bulb. Going from big to small, each question cuts the search in half.
Working a cluster outage is the same: kubectl get nodes (does the API answer at all, which nodes are sick?), kubectl get pods -A (what isn't running?), kubectl get events (what just happened?). Then branch: control plane, node, or pod. And keep a log of symptom to cause, written the way you first saw the symptom, because that's how you'll recognise it next time.
Top-down, every time
The problem. In a real outage nobody tells you which part is broken, and the first thing you look at is often a symptom of something else. A fixed order of questions gets you to the broken layer in a minute instead of an hour.
What you need to know already: everything in this chapter so far - the node tour (18.1), static pods (18.3), certificates (18.11), node lifecycle (18.25) and the failure catalogue (18.28, 18.34).
When "the cluster is broken", the question is which layer. Three commands tell you in under a minute: Three commands tell
$ kubectl get nodes # 1. does the apiserver answer at all? which nodes are sick?
$ kubectl get pods -A | grep -v Running # 2. what is not running, and where?
$ kubectl get events -A --sort-by=.lastTimestamp | tail -20 # 3. what just happened?
Then branch:
kubectl cannot connect - the control plane. Everything else is noise until this is fixed.
connection refused -> apiserver not running: ssh cp-1; sudo crictl ps -a --name kube-apiserver
gone from crictl -> manifest does not parse (journalctl -u kubelet | grep manifest)
Exited over and over -> crictl logs <id>: bad flag, missing file, etcd unreachable
kubelet itself down on cp-1 -> systemctl status kubelet
x509: certificate has expired -> kubeadm certs check-expiration, renew, restart the static pods
Unauthorized -> your kubeconfig's client cert (copy the current admin.conf)
etcdserver: request timed out -> etcd: crictl ps --name etcd, crictl logs, disk full on cp-1
a node is NotReady - describe node Conditions, then on the node: systemctl status kubelet, journalctl -u kubelet -n 30, systemctl status containerd, ls /etc/cni/net.d, df -h /.
pods are sick, nodes are fine - the catalogue (18.28, 18.34): describe pod Events, logs --previous.
Two rules that save hours:
- Read the whole message. kubelet and kubeadm errors name the file, flag or path.
"command failed" err="... /var/lib/kubelet/config.yaml ..." is not "kubelet broken", it is "that file". - Fix the cause, not the symptom. Restarting the kubelet on a node with swap on "fixes" it until the next reboot. Deleting Evicted pods on a full disk makes new ones. Ask "what made it that way" before "how do I make it green".
Prove the fix
A fix is done when you have seen the symptom go away, not the cause: kubectl get nodes all Ready, the workload Ready and stable (no restarts for a minute), and, for anything that survives reboots (swap, enabled units, sysctl), the persistent file checked. The capstone incidents check exactly that.
The log
The plan's exercise: break the cluster on purpose and keep a log of symptom -> cause for each break. The value is in writing the symptom the way you first saw it ("kubectl: connection refused", not "apiserver manifest broken"), because that is what you will see next time. The last mission of this series has you write yours.
What you can now do
- Place any cluster problem in its layer (control plane, node, workload) with three commands.
- Follow the branch for that layer to the file, flag or unit that is wrong.
- Prove a fix, including across a reboot.
Why it helps
In a real incident, and in the admin exam's troubleshooting tasks, time goes to wandering: checking pods when the API server is down, restarting kubelets on healthy nodes, reading documentation for an error that already named the broken file. A fixed top-down order gets you to the right layer in under a minute.
The "fix the cause, prove the fix" rules are what separate a good on-call engineer from a lucky one: restarting a kubelet on a node with swap turned back on "fixes" it until the next reboot; a fix is only done when the symptom is gone and the persistent file is checked. And the symptom-to-cause log you build here becomes your personal runbook, the thing you'll draw on in interviews when asked "tell me about an incident you debugged".
FAQ
Why start with kubectl get nodes?
Because it answers two questions at once. If it fails, the API server is unreachable and nothing else you do through kubectl matters: go to the control plane. If it works, it immediately shows which nodes are sick, which tells you whether the problem is one machine, several, or none (and then it's in the workloads).
What does "read the whole message" mean in practice?
Kubernetes errors are usually specific. A kubelet error like "command failed" err="failed to load kubelet config file, path: /var/lib/kubelet/config.yaml" isn't "the kubelet is broken", it's "that file". A FailedScheduling message lists a reason per node. An image pull error quotes the registry. The fix is often in the last clause of the line people skim.
When is a fix actually done?
When you've seen the symptom go away, not just the cause you changed: nodes Ready, the workload Ready with no restarts for a while, the request that failed now succeeding. And for anything that must survive a reboot (swap in fstab, an enabled unit, a sysctl file), you've checked the persistent configuration, not just the running state.
Why write the symptom the way I first saw it?
Because next time you'll see the symptom, not the cause. "kubectl: connection refused" is what you'll search your notes for at 3am; "apiserver manifest broken" is only useful once you already know. A log of symptom to cause to fix turns past incidents into fast pattern matching.
What if several things look broken at once?
Fix from the bottom of the dependency chain up. The API server and etcd first, because without them nothing else can be observed or changed; then nodes (kubelet, runtime, CNI); then cluster add-ons like CoreDNS; then workloads. Many "broken" pods are just downstream of one lower-layer failure and recover on their own.
In an interview Mid
What are your first commands when someone says "the cluster is broken"?
Find the layer first - three commands:
kubectl get nodes - does the API server answer at all, and which nodes are sick?kubectl get pods -A | grep -v Running - what is not running, and where?kubectl get events -A --sort-by=.lastTimestamp | tail - what just happened?
Then branch:
- kubectl cannot connect - the control plane:
ssh cp-1, sudo crictl ps -a --name kube-apiserver, crictl logs, the kubelet journal for manifest errors, kubeadm certs check-expiration for x509. Everything else is noise until this works. - A node NotReady -
describe node Conditions, then on the node: kubelet, containerd, /etc/cni/net.d, df -h /. - Pods sick, nodes fine - the workload:
describe pod Events, logs --previous.
Two rules: read the whole error - it usually names the file or flag; and fix the cause, not the symptom (restarting a kubelet with swap on works only until the reboot). A fix is done when the symptom is gone and the persistent setting is checked.
Also asked: How do you decide whether a problem is in the control plane, a node, or the workload? · How would you approach an incident where many services in a cluster fail at once? · How do you prove a fix will survive a reboot?