Why know the control plane
When a deploy "succeeds" but nothing runs, or a pod sits in Pending for ever, the question is always which part of the brain stopped doing its job. The control plane is only four programs, each with one job. Know the four jobs and every "stuck" symptom points at one of them.
What you need to know already: the cluster, the API server and kubectl (15.1), objects, kinds and -o yaml (15.3), the HTTPS request/response cycle (9.21), JSON and jq-style paths (7.11), grep -E (7.1).
Four components, one rule
+------------------+
kubectl ------> | kube-apiserver | <---- kube-scheduler
kubelets ------> | (the only one | <---- kube-controller-manager
kube-proxy ----> | that talks to | <---- cloud-controller-manager
| etcd) |
+--------+---------+
|
+--+---+
| etcd |
+------+
| component | its one job |
|---|---|
| kube-apiserver | the front door: checks and stores every change, answers every question |
| etcd | the database where every object is stored |
| kube-scheduler | picks a node for each new pod |
| kube-controller-manager | runs the controllers: small programs that make reality match what was asked for |
The rule: only the API server reads and writes etcd. Every other component - the scheduler, the controllers, the agent on every node (the kubelet, 15.7), kube-proxy (15.7), you - is a client of the API server. They watch (keep a connection open and get told about every change to the objects they care about) and write results back. Nothing talks to anything else directly.
On your cluster they all run on cp-1. k get pods -n kube-system -o wide lists the pods in the kube-system namespace (where Kubernetes keeps its own parts), with node and IP columns; grep cp-1 keeps the cp-1 lines:
$ k get pods -n kube-system -o wide | grep cp-1
etcd-cp-1 1/1 Running 0 12d 10.64.0.10 cp-1 <none> <none>
kube-apiserver-cp-1 1/1 Running 0 12d 10.64.0.10 cp-1 <none> <none>
kube-controller-manager-cp-1 1/1 Running 0 12d 10.64.0.10 cp-1 <none> <none>
kube-scheduler-cp-1 1/1 Running 0 12d 10.64.0.10 cp-1 <none> <none>
(The columns are NAME, READY, STATUS, RESTARTS, AGE, IP, NODE and two you can ignore; 15.14 reads them properly.) Their IP is the node's IP (10.64.0.10), not a pod IP from 10.244.x.x. They run with hostNetwork: true - directly on the node's network, like docker run --network host (11.15). They have to: they start before the pod network exists.
kube-apiserver
The front door. Every request - from you, a controller or a node - goes through the same steps, in this order:
1. authentication who are you? client cert, token...
2. authorization may you do this? the permission rules
3. admission mutating, then validating plugins
4. validation is the object well-formed? (schema, field rules)
5. persist write to etcd, bump resourceVersion, notify watchers
- Authentication - checking identity. Your kubeconfig's client certificate (15.1) says "I am kubernetes-admin".
- Authorization - checking permission: is kubernetes-admin allowed to create pods here? (Yes: it is cluster-admin.)
- Admission - plugins that may change the object on its way in (mutating) and then reject it (validating). Organisations put rules here: "no containers running as root", "images only from our registry".
- Validation - checking the object against its schema, the list of fields its kind allows and their types (what
k explainshows). - Persist - the object is saved in etcd with a new resourceVersion (a counter that changes on every write) and everyone watching is told.
Later (Ch 17): authorization is RBAC - roles that grant who may do what; you write them there.
You can see pieces of this in the API server's own settings. -o yaml prints its pod object; grep -E keeps lines matching any of the flags; -- tells grep that the pattern starting with - is not an option:
$ k get pod kube-apiserver-cp-1 -n kube-system -o yaml | grep -E -- '--(authorization|enable-admission|etcd-servers|service-cluster)'
- --authorization-mode=Node,RBAC
- --enable-admission-plugins=NodeRestriction
- --etcd-servers=https://127.0.0.1:2379
- --service-cluster-ip-range=10.96.0.0/12
Line by line: permissions are checked by two authorizers (Node for node agents, RBAC for everyone else); one extra admission plugin is on; etcd is reached at 127.0.0.1:2379 (same machine, port 2379); Services get IPs from 10.96.0.0/12 (15.1).
Admission changes objects silently. Create a pod and read it back with -o yaml: it has a volume called kube-api-access-xxxxx and two tolerations you never wrote. Built-in admission plugins added them (the volume gives the pod a token to call the API; tolerations come up in 15.7).
Validation errors come back to kubectl word for word:
# bad.yaml: a Deployment whose selector says app=api but whose template says app=web
k apply -f bad.yaml
The Deployment "bad" is invalid: spec.template.metadata.labels: Invalid value: map[string]string{"app":"web"}: `selector` does not match template `labels`
(A Deployment finds its pods by a label, app=api; here its own pod template says app=web, so it could never find them. 15.26 explains labels.)
# typo.yaml: `replica: 3` instead of `replicas: 3`
k apply -f typo.yaml
Error from server (BadRequest): error when creating "typo.yaml": Deployment in version "v1" cannot be handled as a Deployment: strict decoding error: unknown field "spec.replica"
That second one is strict field validation (the default since v1.27): an unknown field is an error, not silently dropped. Before 1.27, replica: 3 was ignored and you got one replica and a long afternoon.
etcd
etcd is a small, very reliable key-value database (like a dictionary: key -> value). Every object in the cluster is one key under /registry/:
/registry/pods/default/web-7d9f5c-abcde
/registry/deployments/default/web
/registry/services/specs/default/web
Production runs etcd on several machines that keep identical copies. They agree on every write with a voting protocol called Raft: a write counts only when a majority - the quorum - of members have stored it. Quorum is floor(n/2) + 1 (more than half):
members quorum failures tolerated
1 1 0
2 2 0 <- worse than 1: twice the hardware, same tolerance
3 2 1
4 3 1 <- no better than 3
5 3 2
That table is the whole answer to "why is etcd always an odd number": an extra even member is one more machine that can fail, without one more failure you can survive. Production control planes run 3 (or 5). This lab runs 1, which is why it is a lab.
If etcd loses quorum the cluster keeps running - containers do not stop - but nothing can change: no scheduling, no scaling, no deletes; every write errors. That is why etcd gets backed up (etcdctl snapshot save, a later chapter's job).
kube-scheduler
The scheduler watches for pods that have no node yet (their spec.nodeName field is empty), and for each one:
- Filter - which nodes could run it: enough free CPU and memory for what the pod says it needs (its requests - the CPU and memory a pod reserves on a node), the pod's node rules match, the node's "keep out" marks (taints, 15.7) are allowed, ports are free.
- Score - rank the nodes that passed (spread copies of the same app apart, prefer less loaded nodes) and pick the highest.
- Bind - write a Binding: "this pod goes to that node". That sets
spec.nodeName, and that is all the scheduler writes. It never starts anything.
When no node passes the filter, you get the scheduler's most useful message, as an Event (a log line attached to the pod, 15.11):
Warning FailedScheduling 5s default-scheduler 0/3 nodes are available: 1 node(s) had untolerated taint {node-role.kubernetes.io/control-plane: }, 2 Insufficient memory. preemption: 0/3 nodes are available: 1 Preemption is not helpful for scheduling, 2 No preemption victims found for incoming pod.
Read it as a tally per node: 0/3 nodes are available - none of 3 fit. 1 node(s) had untolerated taint {...control-plane} - cp-1 was excluded by its keep-out mark. 2 Insufficient memory - both workers lacked the memory the pod asked for. The preemption: part says evicting smaller pods would not help.
kube-controller-manager
A controller is a loop that watches one kind of object and keeps making reality match it (15.9 shows the loop in action). The kube-controller-manager is one program running dozens of them:
deployment controller Deployment -> ReplicaSets
replicaset controller ReplicaSet -> Pods
statefulset, daemonset, job, cronjob controllers
node controller notices dead kubelets, marks nodes NotReady, evicts
endpoints / endpointslice controllers Service -> the IPs of its ready pods
serviceaccount controller a "default" ServiceAccount in every namespace
namespace controller deletes everything inside a namespace being deleted
garbage collector deletes objects whose owners are gone
Most names are kinds you meet later in this chapter (see the table in 15.3). The pattern is what matters: one controller per kind, each owning one job.
Its settings, extracted with jsonpath (a way to pick one field out of an object, like a jq path - 15.38 teaches it), then tr ',' '\n' puts each flag on its own line:
$ k get pod kube-controller-manager-cp-1 -n kube-system -o jsonpath='{.spec.containers[0].command}' | tr ',' '\n' | grep -E 'controllers|cluster-cidr|leader'
"--cluster-cidr=10.244.0.0/16"
"--controllers=*,bootstrapsigner,tokencleaner"
"--leader-elect=true"
--cluster-cidr is the pod network; --controllers=*,... = run all the default controllers plus two extras. --leader-elect=true matters with 3 control-plane nodes: all three run a controller-manager, but only one - the leader, the holder of a lock object called a Lease - acts. If it dies, another takes the lock. The scheduler works the same way.
cloud-controller-manager
When a cluster runs on a cloud provider, one more component talks to the cloud's own API: it creates a cloud load balancer (9.23) when you ask for a public Service, and deletes Node objects for VMs that no longer exist. On a managed cloud cluster the provider runs it for you. On this bare kubeadm cluster there is none - so a Service that asks for a cloud load balancer stays <pending> for ever here.
How the control plane runs itself: static pods
Chicken and egg: pods are created through the API server, but the API server is a pod. The answer is static pods: the kubelet (the node agent) on cp-1 reads YAML files from a directory, /etc/kubernetes/manifests/, and runs them directly - no API server involved:
/etc/kubernetes/manifests/
etcd.yaml
kube-apiserver.yaml
kube-controller-manager.yaml
kube-scheduler.yaml
The kubelet then creates a read-only copy in the API - a mirror pod - so you can see them. kube-apiserver-cp-1 is one. Recognise a mirror pod by its owner and a marker annotation (a note on the object, 15.26):
$ k get pod etcd-cp-1 -n kube-system -o jsonpath='{.metadata.ownerReferences[0].kind}/{.metadata.ownerReferences[0].name}{"\n"}'
Node/cp-1
$ k get pod etcd-cp-1 -n kube-system -o jsonpath='{.metadata.annotations.kubernetes\.io/config\.source}{"\n"}'
file
- the name ends in
-<nodename>(etcd-cp-1) - the owner (
ownerReferences, 15.9) is the Node, not a controller - the annotation
kubernetes.io/config.sourcesaysfile
Deleting a mirror pod does nothing lasting - the kubelet recreates it from the file. To change the API server you edit the manifest file on cp-1 and the kubelet restarts it. Break the file and the API server disappears - the kubelet is only following the file, the same way systemd follows a unit file (2.1).
Where the rest of kube-system comes from
Not everything in kube-system is static. k get deploy,ds -n kube-system lists the Deployments and DaemonSets there (15.3's table):
$ k get deploy,ds -n kube-system
NAME READY UP-TO-DATE AVAILABLE AGE
deployment.apps/calico-kube-controllers 1/1 1 1 12d
deployment.apps/coredns 2/2 2 2 12d
deployment.apps/metrics-server 1/1 1 1 12d
NAME DESIRED CURRENT READY UP-TO-DATE AVAILABLE NODE SELECTOR AGE
daemonset.apps/calico-node 3 3 3 3 3 kubernetes.io/os=linux 12d
daemonset.apps/kube-proxy 3 3 3 3 3 kubernetes.io/os=linux 12d
CoreDNS (the cluster DNS server) is an ordinary Deployment with 2 copies (2/2 ready). kube-proxy and calico-node are DaemonSets - one pod per node, including cp-1 (DESIRED 3). metrics-server collects CPU/memory usage for kubectl top. These are add-ons: they run on the cluster rather than making it.
What you can now do:
- name the four control-plane components and the one job of each
- say why only the API server touches etcd, and why etcd has 3 or 5 members
- recognise a static (mirror) pod and know it is changed by editing a file