OnCallReady

Chapter 18 Kubernetes: Cluster Operations & Troubleshooting

The nodes as machines: ssh, the kubelet as a systemd service, crictl, static pods and the control-plane bootstrap; kubeadm join, certificates and renewal, upgrades and version skew, etcd backup and restore, node lifecycle, and the full failure catalogue - broken on purpose and fixed from the node up.

In plain words

Imagine being the caretaker of a school, not a teacher. Teachers run the lessons (that's the workloads). You look after the building itself: the boiler, the fuse box, the keys, the register of who's enrolled. When the boiler breaks, nobody can teach, and the fix isn't in a classroom; it's in the basement with a torch. You also do the planned jobs: replacing the boiler once a year, backing up the enrolment register, and closing a wing for repairs without stranding anyone.

That's cluster operations. The nodes are Linux machines with a kubelet (a systemd service) and containerd. The control plane runs as static pods from /etc/kubernetes/manifests. kubeadm builds it, issues its certificates and upgrades it. etcd is the register. When kubectl stops answering, you ssh in and use systemctl, journalctl and crictl.

Why it matters on call

This is the chapter for when Kubernetes itself is broken, which is exactly when a platform engineer earns their pay. "kubectl: connection refused", every node NotReady after exactly one year, a node stuck in DiskPressure evicting pods, an upgrade that has to happen this quarter because the minor version goes out of support. Each needs you to drop below kubectl onto the machines, and your Block 1 skills (systemd, journald, the filesystem) are what you use there.

It's also the core of the admin exam's troubleshooting domain (the largest, around 30%) and cluster architecture tasks: upgrade a kubeadm cluster one minor, back up and restore etcd, fix a broken node or control-plane component. On a managed cloud cluster you won't run kubeadm, but the same concepts (skew, drains, certificates, node conditions) show up in every node pool upgrade and incident.

Lessons

  1. Under kubectl: the nodes are Linux machines
  2. Static pods: how the control plane runs itself
  3. kubeadm init and join: what they actually do
  4. kubeadm's certificates: who trusts whom, and for how long
  5. Version skew: what may run with what
  6. The kubeadm upgrade, step by step
  7. etcd: the cluster in one database, and how to back it up
  8. Restoring etcd from a snapshot
  9. Node lifecycle: cordon, drain, NotReady and taint-based eviction
  10. The failure catalogue I: pods that will not run
  11. The failure catalogue II: nodes, Services, DNS, certificates
  12. Working a cluster outage without looking anything up

49 hands-on labs (missions, incidents and drills) run in the terminal: Open this chapter in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.

Questions people ask

If I use a managed cloud cluster, do I still need to know kubeadm?

For the admin exam, yes: the exam clusters are kubeadm-based. For the job, the specific commands matter less, but the model matters a lot: nodes are Linux VMs running kubelet and containerd, control-plane components have versions and certificates, upgrades go control plane first then nodes, drains respect PDBs. A managed service automates the steps; when one fails, you debug with the same concepts.

Why does the kubelet run as a systemd service and not as a pod?

Because it's what runs pods. It has to exist before any pod can start, including the control plane's own static pods, so systemd starts it at boot. That means everything from the systemd chapter applies: systemctl status kubelet, journalctl -u kubelet, drop-ins in kubelet.service.d, Restart=always.

What's the difference between kubectl and crictl?

kubectl asks the API server what it believes should be running. crictl asks the container runtime on one node what's actually running there. crictl works when the API server is down, which is exactly when you need it: sudo crictl ps -a to see crashed control-plane containers and sudo crictl logs <id> to read why.

How often do I need to upgrade a cluster?

Kubernetes ships a minor release roughly every four months, and each minor gets about a year of patches. Staying supported means roughly three minor upgrades a year, one minor at a time. On kubeadm, upgrading on schedule also renews the one-year leaf certificates, so an upgraded cluster never hits certificate expiry.

Is an etcd backup the same as backing up my apps?

No. An etcd snapshot holds the cluster's state: every object, including Secrets. It doesn't hold data in PersistentVolumes. Restoring it rewinds the whole cluster to that moment. App data needs its own backups (volume snapshots, a backup tool such as Velero, database tooling), and most apps should also be redeployable from Git without etcd.