When kubectl stops being enough
The problem. kubectl answers "connection refused", or a node is NotReady and describe says nothing useful. kubectl only talks to the apiserver; when the apiserver or a node's agent is the thing that is broken, you have to go to the machine itself - with the Linux tools from chapters 1-5.
What you need to know already: ssh (1.1), systemd units, drop-ins, Restart= and the start limit (2.1-2.12), journalctl (2.30), df (4.19), the node components - kubelet, containerd, CNI, kube-proxy (15.7), the control plane (15.5).
The lessons use chapter 15's speed kit. If this shell has no k (you jumped in here), define it now:
$ alias k=kubectl
Everything so far went through the apiserver: kubectl get, describe, logs. That works as long as the control plane is healthy and the node's kubelet can report. Cluster operations is the job you do when that is not true: the apiserver is down, a node is NotReady, certificates expired. Then the only way in is the one you used in Block 1: ssh to the machine and read systemd, the journal, the files.
This lab cluster is three kubeadm VMs next to oncall-lab (kubeadm = the official tool that turns plain Linux machines into a Kubernetes cluster; lesson 18.6 shows what it does):
| node | IP | role | what runs there |
|---|---|---|---|
| cp-1 | 10.64.0.10 | control plane | kubelet, containerd, static pods (pods the kubelet runs from files, lesson 18.3): etcd, kube-apiserver, kube-controller-manager, kube-scheduler; plus kube-proxy, calico-node (the network plugin) |
| worker-1 | 10.64.0.11 | worker | kubelet, containerd, kube-proxy, calico-node, your workloads |
| worker-2 | 10.64.0.12 | worker | same |
(simulator) Your key is already in authorized_keys on every node and the names resolve from oncall-lab, so ssh cp-1 just works. On real VMs that is the first thing you set up. The prompt tells you where you are - it changes to learner@cp-1:~$. exit (or Ctrl+D) brings you back.
learner@oncall-lab:~$ ssh worker-1
Welcome to Ubuntu 24.04.3 LTS (GNU/Linux 6.8.0-85-generic aarch64)
...
Usage of /: 38.7% of 18.59GB Users logged in: 0
Memory usage: 24% IPv4 address for enp0s1: 10.64.0.11
...
learner@worker-1:~$ hostname
worker-1
learner@worker-1:~$ exit
logout
Connection to worker-1 closed.
learner@oncall-lab:~$
For a single command you do not even need a session - ssh runs it and returns:
$ ssh worker-1 systemctl is-active kubelet
active
$ ssh cp-1 'sudo crictl ps --name etcd'
Quote the command when it contains a pipe or $(...): unquoted, your local shell runs that part. ssh cp-1 sudo cat /etc/kubernetes/admin.conf | grep server greps on oncall-lab; ssh cp-1 'sudo cat /etc/kubernetes/admin.conf | grep server' greps on cp-1. Same output here, very different when the thing you pipe into only exists on one side.
The nodes are cloud-image Ubuntu (the ready-made Ubuntu disk image cloud providers use): your user has passwordless sudo (/etc/sudoers.d/90-cloud-init-users). Almost everything below needs it - the kubeadm files are root-only on purpose.
The kubelet is just a systemd service
This is the single most useful fact in this chapter. The kubelet is not a pod. It is a binary (/usr/bin/kubelet) that systemd starts, so everything from the systemd chapter applies:
Log in to worker-1 for the next part (the prompts say which machine answers):
$ ssh worker-1
learner@worker-1:~$ systemctl status kubelet
● kubelet.service - kubelet: The Kubernetes Node Agent
Loaded: loaded (/usr/lib/systemd/system/kubelet.service; enabled; preset: enabled)
Drop-In: /usr/lib/systemd/system/kubelet.service.d
└─10-kubeadm.conf
Active: active (running) since Thu 2026-09-10 16:47:03 UTC; 1 week 5 days ago
Docs: https://kubernetes.io/docs/
Main PID: 844 (kubelet)
Tasks: 14 (limit: 4598)
Memory: 79.3M (peak: 83.0M)
CPU: 48min 32.779s
CGroup: /system.slice/kubelet.service
└─844 /usr/bin/kubelet --bootstrap-kubeconfig=/etc/kubernetes/bootstrap-kubelet.conf --kubeconfig=/etc/kubernetes/kubelet.conf --config=/var/lib/kubelet/config.yaml --container-runtime-endpoint=unix:///var/run/containerd/containerd.sock --node-ip=10.64.0.11
Read it the way you read any unit:
- Loaded - the unit file is the package's, in
/usr/lib/systemd/system, andenabledmeans it starts at boot. A kubelet that isdisabledworks until the first reboot. That is a classic. - Drop-In
10-kubeadm.conf- this is where the command line actually comes from. The base unit only saysExecStart=/usr/bin/kubelet; the drop-in clears it and rebuilds it from environment variables:
$ systemctl cat kubelet
# /usr/lib/systemd/system/kubelet.service
[Unit]
Description=kubelet: The Kubernetes Node Agent
...
[Service]
ExecStart=/usr/bin/kubelet
Restart=always
StartLimitInterval=0
RestartSec=10
...
# /usr/lib/systemd/system/kubelet.service.d/10-kubeadm.conf
[Service]
Environment="KUBELET_KUBECONFIG_ARGS=--bootstrap-kubeconfig=/etc/kubernetes/bootstrap-kubelet.conf --kubeconfig=/etc/kubernetes/kubelet.conf"
Environment="KUBELET_CONFIG_ARGS=--config=/var/lib/kubelet/config.yaml"
EnvironmentFile=-/var/lib/kubelet/kubeadm-flags.env
EnvironmentFile=-/etc/default/kubelet
ExecStart=
ExecStart=/usr/bin/kubelet $KUBELET_KUBECONFIG_ARGS $KUBELET_CONFIG_ARGS $KUBELET_KUBEADM_ARGS $KUBELET_EXTRA_ARGS
Restart=always+RestartSec=10+StartLimitInterval=0: a kubelet that crashes is restarted every 10 seconds forever (the start limit is disabled - remember why that matters from chapter 2). So a broken kubelet does not show asfailed; it shows asactivating (auto-restart), over and over.
Where kubeadm puts things
Memorise this map; every troubleshooting path goes through it.
| path | what |
|---|---|
/var/lib/kubelet/config.yaml | the kubelet's configuration (KubeletConfiguration): cgroupDriver (who manages cgroups - must be systemd), clusterDNS, staticPodPath, eviction thresholds |
/var/lib/kubelet/kubeadm-flags.env | the few flags kubeadm passes on the command line: runtime endpoint, node IP |
/etc/kubernetes/kubelet.conf | the kubelet's kubeconfig: which apiserver, which client certificate |
/var/lib/kubelet/pki/ | the kubelet's own certificates (kubelet-client-current.pem rotates itself) |
/etc/kubernetes/manifests/ | static pod manifests (only cp-1 has files here) |
/etc/kubernetes/pki/ | cluster CA and every control-plane certificate (cp-1) |
/etc/kubernetes/admin.conf | the cluster-admin kubeconfig (cp-1, root only) |
/var/lib/etcd/ | etcd's data - the whole cluster state (cp-1) |
/etc/cni/net.d/ | the CNI config the runtime uses to give pods an IP (written by calico-node) |
/opt/cni/bin/ | the CNI plugin binaries |
/etc/containerd/config.toml | containerd config (SystemdCgroup = true must match the kubelet's cgroupDriver: systemd) |
The kubelet journal
When a node misbehaves, the first two commands are the same as for any service:
$ systemctl status kubelet
$ journalctl -u kubelet -n 50 --no-pager
Kubelet lines are klog lines (klog = the logging library all Kubernetes components use). Learn to read one:
Sep 22 20:00:19 worker-1 kubelet[844]: E0922 20:00:19.000261 844 kubelet.go:3117] "Container runtime network not ready" networkReady="NetworkReady=false reason:NetworkPluginNotReady message:Network plugin returns error: cni plugin not initialized"
Sep 22 20:00:19 worker-1 kubelet[844]:- journald's prefix: time, host, process.E0922 20:00:19.000261- klog's own: severity letter (I info, W warning, E error, F fatal), month+day, time with microseconds.844- the thread/process id,kubelet.go:3117]- the source file and line.- then a quoted message and
key="value"pairs. The message is the what; theerr=field is almost always the why.
Useful filters (all the journald flags from chapter 2 work - -p err = priority error and worse, --since, -g = grep the message, -f = follow):
$ journalctl -u kubelet -p err --since "10 min ago" # errors only, recent
$ journalctl -u kubelet -g "command failed" # grep the messages
$ journalctl -u kubelet -f # follow while you fix
$ journalctl -u containerd -n 20 # the runtime's side
The four kubelet lines you will see most, and what they mean:
"command failed" err="failed to load kubelet config file, path: /var/lib/kubelet/config.yaml ..."
The kubelet cannot start at all: config missing (node never joined) or broken. Status shows activating (auto-restart).
"Container runtime network not ready" ... "cni plugin not initialized"
The kubelet runs, but there is no CNI config in /etc/cni/net.d - node NotReady, new pods stuck in ContainerCreating.
"Error updating node status, will retry" err="... dial tcp 10.64.0.10:6443: connect: connection refused"
This node is fine; the apiserver is not. Go look at cp-1.
"Skipping pod synchronization" err="[container runtime is down, PLEG is not healthy ...]"
containerd is down or hung (PLEG = the kubelet's Pod Lifecycle Event Generator, the loop that asks the runtime what changed). systemctl status containerd.
crictl: what the runtime is really running
kubectl get pods asks the apiserver what it believes. crictl (the command-line client for any CRI runtime - CRI = Container Runtime Interface, the API the kubelet uses to talk to containerd, 15.7) asks the container runtime on this node what is actually running - and it works when the apiserver is dead, which is exactly when you need it.
From worker-1 over to cp-1:
$ exit
$ ssh cp-1
learner@cp-1:~$ sudo crictl ps
CONTAINER IMAGE CREATED STATE NAME ATTEMPT POD ID POD NAMESPACE
1902945fdc185 ccff97a6bff7c 12 days ago Running calico-node 0 ed695692a270b calico-node-scdpx kube-system
43bbd666b1280 a64cbce4cbc68 12 days ago Running kube-proxy 0 7cce480cb917b kube-proxy-wsrxs kube-system
96899c9e92697 fe9a2fec5feea 12 days ago Running etcd 0 f32a9b52adc65 etcd-cp-1 kube-system
00295ffb3e0cc 8cbcefceeb5be 12 days ago Running kube-apiserver 0 dd5be0cbefec8 kube-apiserver-cp-1 kube-system
a9d4b2fb857f6 adcfcbdc5c24a 12 days ago Running kube-controller-manager 0 7c377cc2c51eb kube-controller-manager-cp-1 kube-system
7381cca349927 afc84fdef8fcc 12 days ago Running kube-scheduler 0 a2498f1691fa0 kube-scheduler-cp-1 kube-system
- Without
sudo:FATA[0000] validate service connection: ... dial unix /run/containerd/containerd.sock: connect: permission denied- the socket is root-only. ATTEMPTis the restart count.crictl ps -aalso shows exited containers - for a crash-looping apiserver that is the list of dead attempts.crictl logs <id>prints a container's log straight from the node. Whenkubectl logscannot reach anything, this still works.crictl podslists pod sandboxes,crictl imagesthe images,crictl infothe runtime's own view of its conditions (includingNetworkReady).
The mapping to what you know:
| question | through the apiserver | on the node |
|---|---|---|
| what is running | kubectl get pods -o wide | sudo crictl ps |
| why did it die | kubectl logs --previous | sudo crictl ps -a + sudo crictl logs <id> |
| is the node healthy | kubectl describe node (Conditions) | systemctl status kubelet, journalctl -u kubelet |
| disk | describe node DiskPressure | df -h /, du |
Back to oncall-lab:
$ exit
The crictl ps columns: CONTAINER (short container id), IMAGE (image id), CREATED, STATE, NAME (the container name from the pod spec), ATTEMPT (restarts), POD ID, POD (the pod name), NAMESPACE.
The rule for the rest of the chapter
Start from the apiserver, drop to the node when the apiserver cannot answer or the answer points at a node. kubectl get nodes says worker-2 is NotReady - you go to worker-2. kubectl itself says "connection refused" - you go to cp-1 and ask crictl what happened to the apiserver.
What you can now do
- ssh to a node, read the kubelet unit, its drop-in and its journal.
- Find the kubeadm files on a node from the map above.
- Use
crictl ps,ps -aandlogswhen kubectl cannot help.