OnCallReady

Lesson 18.3 · Kubernetes: Cluster Operations & Troubleshooting · 19 min read

Static pods: how the control plane runs itself

In plain words

Imagine a factory whose manager (the API server) tells the machines what to build. But who switches the manager's own desk lamp on in the morning? The night watchman has a note taped to the wall: "at dawn, turn on these four things". He doesn't need anyone's permission; the note is enough. Take the note down and he turns them off.

That note is /etc/kubernetes/manifests. The kubelet runs every pod file it finds there, with no scheduler and no API server, which is how etcd, the API server, the controller-manager and the scheduler start before a cluster exists. These are static pods. The kubelet shows them in kubectl as mirror pods (kube-apiserver-cp-1), but deleting a mirror pod changes nothing: you change a static pod by editing its file.

The chicken and the egg

The problem. The control plane itself runs as pods - but pods need a running control plane. Understanding how that loop is broken tells you how to restart, reconfigure and repair the apiserver, scheduler and etcd when kubectl is useless.

What you need to know already: the node tour (18.1), pod manifests (15.14), the control plane components (15.5), mv/cp (1.5), sed -i (7.6).

The apiserver stores pods. The scheduler places pods. The kubelet runs pods it reads from the apiserver. So who runs the apiserver pod, before there is an apiserver?

The kubelet, from files. Besides watching the apiserver, the kubelet watches a directory - staticPodPath in its config, /etc/kubernetes/manifests on kubeadm nodes - and runs every pod manifest it finds there. No scheduler, no apiserver, no controller: file present = pod running, file gone = pod stopped.

On cp-1 (kubectl works there too; exit when you leave the lesson):

$ ssh cp-1
learner@cp-1:~$ sudo ls -l /etc/kubernetes/manifests
total 16
-rw------- 1 root root 2514 Sep 10 16:43 etcd.yaml
-rw------- 1 root root 4062 Sep 10 16:43 kube-apiserver.yaml
-rw------- 1 root root 3389 Sep 10 16:43 kube-controller-manager.yaml
-rw------- 1 root root 1463 Sep 10 16:43 kube-scheduler.yaml

That is the bootstrap order kubeadm relies on:

  1. systemd starts containerd, then the kubelet.
  2. The kubelet reads /etc/kubernetes/manifests and starts etcd and the apiserver as containers - no cluster needed yet.
  3. The apiserver connects to etcd on https://127.0.0.1:2379.
  4. The kubelet can now register the node with the apiserver it just started.
  5. controller-manager and scheduler (also static pods) connect to the apiserver with their kubeconfigs (/etc/kubernetes/controller-manager.conf, scheduler.conf).
  6. kubeadm (or you) apply the add-ons as normal objects: CoreDNS Deployment, kube-proxy DaemonSet, the CNI (calico here). Those go through the scheduler like any workload.

Mirror pods

The kubelet also creates a mirror pod (a read-only copy that exists only so kubectl can see it) in the apiserver for every static pod, so kubectl can see it. The name is <pod name>-<node name>:

$ kubectl get pods -n kube-system -o wide | grep cp-1
etcd-cp-1                      1/1   Running   0   12d   10.64.0.10   cp-1
kube-apiserver-cp-1            1/1   Running   0   12d   10.64.0.10   cp-1
kube-controller-manager-cp-1   1/1   Running   0   12d   10.64.0.10   cp-1
kube-scheduler-cp-1            1/1   Running   0   12d   10.64.0.10   cp-1

A mirror pod is read-only as far as the real pod goes:

$ kubectl delete pod -n kube-system kube-scheduler-cp-1
pod "kube-scheduler-cp-1" deleted
$ kubectl get pod -n kube-system kube-scheduler-cp-1
NAME                  READY   STATUS    RESTARTS   AGE
kube-scheduler-cp-1   1/1     Running   0          3s

The container never stopped - you deleted the mirror, the kubelet recreated it. You can tell a static pod from its annotations and owner:

$ kubectl get pod -n kube-system kube-scheduler-cp-1 -o jsonpath='{.metadata.ownerReferences[0].kind}{"\n"}{.metadata.annotations.kubernetes\.io/config\.source}{"\n"}'
Node
file

(ownerReferences = the object that owns this one, 15.16; config.source = where the kubelet got the pod from.) Owner Node, source file. A DaemonSet pod would have owner DaemonSet.

Changing a control-plane component

You change the apiserver by editing its manifest. The kubelet notices the file changed (it re-reads the directory every 20s by default, fileCheckFrequency), stops the old container and starts a new one. There is no kubectl apply for static pods, and no kubectl rollout restart.

# an illustration
sudo vi /etc/kubernetes/manifests/kube-apiserver.yaml     # add --v=2, change a flag...
$ sudo crictl ps --name kube-apiserver                      # new CONTAINER id, CREATED seconds ago

(simulator) There is no vi; use sudo nano or sudo sed -i. On the Kubernetes admin exam (Ch 19), vim.

The kubeadm habit: back up before you edit, and back up outside the directory:

$ sudo cp /etc/kubernetes/manifests/kube-apiserver.yaml /root/kube-apiserver.yaml.bak                      # good
$ sudo cp /etc/kubernetes/manifests/kube-apiserver.yaml /etc/kubernetes/manifests/kube-apiserver.yaml.bak   # bad

The second one is a real outage pattern: the kubelet reads every file in the directory (except dotfiles), so it tries to run the backup as a second static pod with the same name.

Restarting one without editing it

To restart a static pod without changing it (for example so it re-reads a renewed certificate), move the manifest out and back:

$ sudo mv /etc/kubernetes/manifests/kube-scheduler.yaml /tmp/
$ sudo crictl ps --name kube-scheduler        # wait until it is gone (up to ~20s)
CONTAINER   IMAGE   CREATED   STATE   NAME   ATTEMPT   POD ID   POD   NAMESPACE
$ sudo mv /tmp/kube-scheduler.yaml /etc/kubernetes/manifests/
$ sudo crictl ps --name kube-scheduler
CONTAINER           IMAGE               CREATED         STATE     NAME             ATTEMPT   POD ID          POD
495f1e8051472       afc84fdef8fcc       4 seconds ago   Running   kube-scheduler   0         aad041f2f34be   kube-scheduler-cp-1

An alternative is sudo crictl stop <container-id>: the kubelet sees the container died and starts it again (ATTEMPT goes up). Moving the file is what the kubeadm docs describe and it works even when crictl is not configured.

What a scheduler outage looks like

While the scheduler is gone, the cluster looks fine - until you create something:

# with kube-scheduler.yaml moved out (the next mission does it for real)
$ kubectl run probe --image=nginx
pod/probe created
$ kubectl get pod probe
NAME    READY   STATUS    RESTARTS   AGE
probe   0/1     Pending   0          40s
$ kubectl describe pod probe | tail -3
Events:            <none>

Pending with no events at all is the signature: a FailedScheduling event is written by the scheduler, so no scheduler, no event. Compare with a normal Pending pod, which always has FailedScheduling explaining why. Same logic for the controller-manager: Deployments stop creating ReplicaSets and ReplicaSets stop creating pods, silently.

A broken manifest

If the YAML does not parse, the kubelet cannot run it and the pod simply disappears - from crictl and from kubectl. The only trace is in the kubelet journal:

# with a kube-apiserver.yaml that does not parse (the incidents build one)
$ journalctl -u kubelet | grep -i manifest
Sep 22 20:03:11 cp-1 kubelet[852]: E0922 20:03:11.000000     852 kubelet.go:2461] "Could not process manifest file" err="/etc/kubernetes/manifests/kube-apiserver.yaml: couldn't parse as pod(yaml: line 38: mapping values are not allowed in this context), please check config file" path="/etc/kubernetes/manifests/kube-apiserver.yaml"

If the YAML parses but the component rejects it (a typo in a flag, a certificate path that does not exist), the container starts, exits, and crash-loops - visible with crictl ps -a and crictl logs:

# an illustration: a flag the apiserver does not know (the incidents build one)
$ sudo crictl ps -a --name kube-apiserver
CONTAINER           IMAGE               CREATED          STATE     NAME             ATTEMPT   POD ID          POD
3a1f09c2d7e88       8cbcefceeb5be       12 seconds ago   Exited    kube-apiserver   4         dd5be0cbefec8   kube-apiserver-cp-1
$ sudo crictl logs 3a1f09c2d7e88
Error: unknown flag: --etcd-server

Two failure shapes, two tools: gone = journal (parse error), crash-looping = crictl logs (the component's own error).

What you can now do

Why it helps

Every control-plane fix on a kubeadm cluster, in real life and on the admin exam, is an edit to one of these four files: add an API server flag, fix an etcd data path, point at a renewed certificate. Knowing the behaviour saves you from the classic mistakes: saving a backup copy inside the manifests directory (the kubelet tries to run it too), expecting kubectl delete pod to restart the scheduler, or panicking when a broken YAML edit makes the API server simply vanish.

It also gives you signatures for incidents: a pod Pending with no events at all means no scheduler is running; a control-plane component gone from crictl means a parse error in the journal; one crash-looping means read crictl logs.

Commands in this lesson

ssh kubectl crictl cp mv journalctl

FAQ

Why did deleting kube-scheduler-cp-1 not restart the scheduler?

You deleted the mirror pod, the read-only reflection the kubelet creates in the API server. The real container, run by the kubelet from its manifest file, never stopped, and the kubelet immediately recreated the mirror. To restart a static pod, move its manifest out of /etc/kubernetes/manifests and back, or crictl stop its container.

Where should I keep a backup of a manifest?

Outside /etc/kubernetes/manifests, for example /root/kube-apiserver.yaml.bak. The kubelet reads every non-hidden file in that directory, so a backup there becomes a second static pod with the same name, which is a real outage pattern. kubeadm itself backs up to /etc/kubernetes/tmp/ during upgrades.

How long after editing a manifest does the change apply?

The kubelet re-checks the directory every 20 seconds by default (fileCheckFrequency). When a file changes, it stops the old container and starts a new one. Watch it with sudo crictl ps --name kube-apiserver: a new container id with a CREATED time of a few seconds means it picked up the change.

How can I tell a static pod from a normal one in kubectl?

Its owner reference is the Node and its kubernetes.io/config.source annotation is file. Its name ends with the node name, like etcd-cp-1. A DaemonSet pod, by contrast, is owned by a DaemonSet and goes through the scheduler (well, the DaemonSet controller), not the kubelet's manifest directory.

The API server disappeared after my edit. What happened?

The YAML probably doesn't parse, so the kubelet can't create the pod at all and it vanishes from both crictl and kubectl. The only trace is in the kubelet's journal: journalctl -u kubelet | grep -i manifest shows "Could not process manifest file" with the line number. If it parses but a flag is wrong, you get a crash loop instead, visible in crictl ps -a and crictl logs.

In an interview Junior

What are static pods, and why does kubeadm use them for the control plane?

A static pod is run by the kubelet straight from a manifest file in its staticPodPath - /etc/kubernetes/manifests on kubeadm - with no scheduler, no API server and no controller: file present = pod running, file removed = pod stopped.

kubeadm uses them to break the chicken-and-egg problem: the API server stores pods, but something must run the API server first. So systemd starts containerd and the kubelet, the kubelet starts etcd and kube-apiserver from files, then the controller-manager and scheduler connect to it.

The kubelet also creates a read-only mirror pod (kube-apiserver-cp-1, owner Node) so kubectl can see it; deleting the mirror changes nothing.

To change a component, edit its manifest - the kubelet restarts it within about 20 s. Back up outside the directory (a backup inside it is run as a second pod). A YAML parse error makes the pod vanish (the kubelet journal says why); a bad flag makes it crash-loop (crictl logs).

Also asked: How do you add a flag to the kube-apiserver on a kubeadm cluster? · A pod is Pending and describe shows no events at all. What is wrong? · How do you restart a static pod without changing it?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.