The chicken and the egg
The problem. The control plane itself runs as pods - but pods need a running control plane. Understanding how that loop is broken tells you how to restart, reconfigure and repair the apiserver, scheduler and etcd when kubectl is useless.
What you need to know already: the node tour (18.1), pod manifests (15.14), the control plane components (15.5), mv/cp (1.5), sed -i (7.6).
The apiserver stores pods. The scheduler places pods. The kubelet runs pods it reads from the apiserver. So who runs the apiserver pod, before there is an apiserver?
The kubelet, from files. Besides watching the apiserver, the kubelet watches a directory - staticPodPath in its config, /etc/kubernetes/manifests on kubeadm nodes - and runs every pod manifest it finds there. No scheduler, no apiserver, no controller: file present = pod running, file gone = pod stopped.
On cp-1 (kubectl works there too; exit when you leave the lesson):
$ ssh cp-1
learner@cp-1:~$ sudo ls -l /etc/kubernetes/manifests
total 16
-rw------- 1 root root 2514 Sep 10 16:43 etcd.yaml
-rw------- 1 root root 4062 Sep 10 16:43 kube-apiserver.yaml
-rw------- 1 root root 3389 Sep 10 16:43 kube-controller-manager.yaml
-rw------- 1 root root 1463 Sep 10 16:43 kube-scheduler.yaml
That is the bootstrap order kubeadm relies on:
- systemd starts containerd, then the kubelet.
- The kubelet reads
/etc/kubernetes/manifestsand starts etcd and the apiserver as containers - no cluster needed yet. - The apiserver connects to etcd on
https://127.0.0.1:2379. - The kubelet can now register the node with the apiserver it just started.
- controller-manager and scheduler (also static pods) connect to the apiserver with their kubeconfigs (
/etc/kubernetes/controller-manager.conf,scheduler.conf). - kubeadm (or you) apply the add-ons as normal objects: CoreDNS Deployment, kube-proxy DaemonSet, the CNI (calico here). Those go through the scheduler like any workload.
Mirror pods
The kubelet also creates a mirror pod (a read-only copy that exists only so kubectl can see it) in the apiserver for every static pod, so kubectl can see it. The name is <pod name>-<node name>:
$ kubectl get pods -n kube-system -o wide | grep cp-1
etcd-cp-1 1/1 Running 0 12d 10.64.0.10 cp-1
kube-apiserver-cp-1 1/1 Running 0 12d 10.64.0.10 cp-1
kube-controller-manager-cp-1 1/1 Running 0 12d 10.64.0.10 cp-1
kube-scheduler-cp-1 1/1 Running 0 12d 10.64.0.10 cp-1
A mirror pod is read-only as far as the real pod goes:
$ kubectl delete pod -n kube-system kube-scheduler-cp-1
pod "kube-scheduler-cp-1" deleted
$ kubectl get pod -n kube-system kube-scheduler-cp-1
NAME READY STATUS RESTARTS AGE
kube-scheduler-cp-1 1/1 Running 0 3s
The container never stopped - you deleted the mirror, the kubelet recreated it. You can tell a static pod from its annotations and owner:
$ kubectl get pod -n kube-system kube-scheduler-cp-1 -o jsonpath='{.metadata.ownerReferences[0].kind}{"\n"}{.metadata.annotations.kubernetes\.io/config\.source}{"\n"}'
Node
file
(ownerReferences = the object that owns this one, 15.16; config.source = where the kubelet got the pod from.) Owner Node, source file. A DaemonSet pod would have owner DaemonSet.
Changing a control-plane component
You change the apiserver by editing its manifest. The kubelet notices the file changed (it re-reads the directory every 20s by default, fileCheckFrequency), stops the old container and starts a new one. There is no kubectl apply for static pods, and no kubectl rollout restart.
# an illustration
sudo vi /etc/kubernetes/manifests/kube-apiserver.yaml # add --v=2, change a flag...
$ sudo crictl ps --name kube-apiserver # new CONTAINER id, CREATED seconds ago
(simulator) There is no vi; use sudo nano or sudo sed -i. On the Kubernetes admin exam (Ch 19), vim.
The kubeadm habit: back up before you edit, and back up outside the directory:
$ sudo cp /etc/kubernetes/manifests/kube-apiserver.yaml /root/kube-apiserver.yaml.bak # good
$ sudo cp /etc/kubernetes/manifests/kube-apiserver.yaml /etc/kubernetes/manifests/kube-apiserver.yaml.bak # bad
The second one is a real outage pattern: the kubelet reads every file in the directory (except dotfiles), so it tries to run the backup as a second static pod with the same name.
Restarting one without editing it
To restart a static pod without changing it (for example so it re-reads a renewed certificate), move the manifest out and back:
$ sudo mv /etc/kubernetes/manifests/kube-scheduler.yaml /tmp/
$ sudo crictl ps --name kube-scheduler # wait until it is gone (up to ~20s)
CONTAINER IMAGE CREATED STATE NAME ATTEMPT POD ID POD NAMESPACE
$ sudo mv /tmp/kube-scheduler.yaml /etc/kubernetes/manifests/
$ sudo crictl ps --name kube-scheduler
CONTAINER IMAGE CREATED STATE NAME ATTEMPT POD ID POD
495f1e8051472 afc84fdef8fcc 4 seconds ago Running kube-scheduler 0 aad041f2f34be kube-scheduler-cp-1
An alternative is sudo crictl stop <container-id>: the kubelet sees the container died and starts it again (ATTEMPT goes up). Moving the file is what the kubeadm docs describe and it works even when crictl is not configured.
What a scheduler outage looks like
While the scheduler is gone, the cluster looks fine - until you create something:
# with kube-scheduler.yaml moved out (the next mission does it for real)
$ kubectl run probe --image=nginx
pod/probe created
$ kubectl get pod probe
NAME READY STATUS RESTARTS AGE
probe 0/1 Pending 0 40s
$ kubectl describe pod probe | tail -3
Events: <none>
Pending with no events at all is the signature: a FailedScheduling event is written by the scheduler, so no scheduler, no event. Compare with a normal Pending pod, which always has FailedScheduling explaining why. Same logic for the controller-manager: Deployments stop creating ReplicaSets and ReplicaSets stop creating pods, silently.
A broken manifest
If the YAML does not parse, the kubelet cannot run it and the pod simply disappears - from crictl and from kubectl. The only trace is in the kubelet journal:
# with a kube-apiserver.yaml that does not parse (the incidents build one)
$ journalctl -u kubelet | grep -i manifest
Sep 22 20:03:11 cp-1 kubelet[852]: E0922 20:03:11.000000 852 kubelet.go:2461] "Could not process manifest file" err="/etc/kubernetes/manifests/kube-apiserver.yaml: couldn't parse as pod(yaml: line 38: mapping values are not allowed in this context), please check config file" path="/etc/kubernetes/manifests/kube-apiserver.yaml"
If the YAML parses but the component rejects it (a typo in a flag, a certificate path that does not exist), the container starts, exits, and crash-loops - visible with crictl ps -a and crictl logs:
# an illustration: a flag the apiserver does not know (the incidents build one)
$ sudo crictl ps -a --name kube-apiserver
CONTAINER IMAGE CREATED STATE NAME ATTEMPT POD ID POD
3a1f09c2d7e88 8cbcefceeb5be 12 seconds ago Exited kube-apiserver 4 dd5be0cbefec8 kube-apiserver-cp-1
$ sudo crictl logs 3a1f09c2d7e88
Error: unknown flag: --etcd-server
Two failure shapes, two tools: gone = journal (parse error), crash-looping = crictl logs (the component's own error).
What you can now do
- Explain how the kubelet bootstraps the control plane from
/etc/kubernetes/manifests. - Change or restart a control-plane component by editing or moving its manifest.
- Recognise a missing scheduler (Pending, no events) and a broken manifest (gone vs crash-looping).