OnCallReady

Lesson 15.5 · Kubernetes: Architecture & Workloads · 28 min read

The control plane: who does what

In plain words

Think of a school office. There is one secretary at the front desk (kube-apiserver) who is the only person allowed to write in the big register (etcd). Everyone else, the timetable planner (kube-scheduler) and the various helpers who make sure things happen (controllers in kube-controller-manager), asks the secretary to look things up and to write things down. The timetable planner only writes "this pupil goes to room 3"; she never walks anyone there.

That's the control plane. The apiserver runs every request through authentication, authorisation, admission and validation before writing to etcd. etcd uses Raft and needs a majority: 3 members tolerate one failure. The scheduler filters, scores and binds (sets spec.nodeName). Controllers reconcile. On kubeadm they run as static pods from /etc/kubernetes/manifests on cp-1.

Why know the control plane

When a deploy "succeeds" but nothing runs, or a pod sits in Pending for ever, the question is always which part of the brain stopped doing its job. The control plane is only four programs, each with one job. Know the four jobs and every "stuck" symptom points at one of them.

What you need to know already: the cluster, the API server and kubectl (15.1), objects, kinds and -o yaml (15.3), the HTTPS request/response cycle (9.21), JSON and jq-style paths (7.11), grep -E (7.1).

Four components, one rule

                     +------------------+
   kubectl  ------>  |  kube-apiserver  |  <----  kube-scheduler
   kubelets ------>  |   (the only one  |  <----  kube-controller-manager
   kube-proxy ---->  |   that talks to  |  <----  cloud-controller-manager
                     |      etcd)       |
                     +--------+---------+
                              |
                           +--+---+
                           | etcd |
                           +------+
componentits one job
kube-apiserverthe front door: checks and stores every change, answers every question
etcdthe database where every object is stored
kube-schedulerpicks a node for each new pod
kube-controller-managerruns the controllers: small programs that make reality match what was asked for

The rule: only the API server reads and writes etcd. Every other component - the scheduler, the controllers, the agent on every node (the kubelet, 15.7), kube-proxy (15.7), you - is a client of the API server. They watch (keep a connection open and get told about every change to the objects they care about) and write results back. Nothing talks to anything else directly.

On your cluster they all run on cp-1. k get pods -n kube-system -o wide lists the pods in the kube-system namespace (where Kubernetes keeps its own parts), with node and IP columns; grep cp-1 keeps the cp-1 lines:

$ k get pods -n kube-system -o wide | grep cp-1
etcd-cp-1                      1/1   Running   0   12d   10.64.0.10   cp-1   <none>   <none>
kube-apiserver-cp-1            1/1   Running   0   12d   10.64.0.10   cp-1   <none>   <none>
kube-controller-manager-cp-1   1/1   Running   0   12d   10.64.0.10   cp-1   <none>   <none>
kube-scheduler-cp-1            1/1   Running   0   12d   10.64.0.10   cp-1   <none>   <none>

(The columns are NAME, READY, STATUS, RESTARTS, AGE, IP, NODE and two you can ignore; 15.14 reads them properly.) Their IP is the node's IP (10.64.0.10), not a pod IP from 10.244.x.x. They run with hostNetwork: true - directly on the node's network, like docker run --network host (11.15). They have to: they start before the pod network exists.

kube-apiserver

The front door. Every request - from you, a controller or a node - goes through the same steps, in this order:

1. authentication   who are you?          client cert, token...
2. authorization    may you do this?      the permission rules
3. admission        mutating, then validating plugins
4. validation       is the object well-formed? (schema, field rules)
5. persist          write to etcd, bump resourceVersion, notify watchers

Later (Ch 17): authorization is RBAC - roles that grant who may do what; you write them there.

You can see pieces of this in the API server's own settings. -o yaml prints its pod object; grep -E keeps lines matching any of the flags; -- tells grep that the pattern starting with - is not an option:

$ k get pod kube-apiserver-cp-1 -n kube-system -o yaml | grep -E -- '--(authorization|enable-admission|etcd-servers|service-cluster)'
    - --authorization-mode=Node,RBAC
    - --enable-admission-plugins=NodeRestriction
    - --etcd-servers=https://127.0.0.1:2379
    - --service-cluster-ip-range=10.96.0.0/12

Line by line: permissions are checked by two authorizers (Node for node agents, RBAC for everyone else); one extra admission plugin is on; etcd is reached at 127.0.0.1:2379 (same machine, port 2379); Services get IPs from 10.96.0.0/12 (15.1).

Admission changes objects silently. Create a pod and read it back with -o yaml: it has a volume called kube-api-access-xxxxx and two tolerations you never wrote. Built-in admission plugins added them (the volume gives the pod a token to call the API; tolerations come up in 15.7).

Validation errors come back to kubectl word for word:

# bad.yaml: a Deployment whose selector says app=api but whose template says app=web
k apply -f bad.yaml
The Deployment "bad" is invalid: spec.template.metadata.labels: Invalid value: map[string]string{"app":"web"}: `selector` does not match template `labels`

(A Deployment finds its pods by a label, app=api; here its own pod template says app=web, so it could never find them. 15.26 explains labels.)

# typo.yaml: `replica: 3` instead of `replicas: 3`
k apply -f typo.yaml
Error from server (BadRequest): error when creating "typo.yaml": Deployment in version "v1" cannot be handled as a Deployment: strict decoding error: unknown field "spec.replica"

That second one is strict field validation (the default since v1.27): an unknown field is an error, not silently dropped. Before 1.27, replica: 3 was ignored and you got one replica and a long afternoon.

etcd

etcd is a small, very reliable key-value database (like a dictionary: key -> value). Every object in the cluster is one key under /registry/:

/registry/pods/default/web-7d9f5c-abcde
/registry/deployments/default/web
/registry/services/specs/default/web

Production runs etcd on several machines that keep identical copies. They agree on every write with a voting protocol called Raft: a write counts only when a majority - the quorum - of members have stored it. Quorum is floor(n/2) + 1 (more than half):

members   quorum   failures tolerated
   1         1          0
   2         2          0     <- worse than 1: twice the hardware, same tolerance
   3         2          1
   4         3          1     <- no better than 3
   5         3          2

That table is the whole answer to "why is etcd always an odd number": an extra even member is one more machine that can fail, without one more failure you can survive. Production control planes run 3 (or 5). This lab runs 1, which is why it is a lab.

If etcd loses quorum the cluster keeps running - containers do not stop - but nothing can change: no scheduling, no scaling, no deletes; every write errors. That is why etcd gets backed up (etcdctl snapshot save, a later chapter's job).

kube-scheduler

The scheduler watches for pods that have no node yet (their spec.nodeName field is empty), and for each one:

  1. Filter - which nodes could run it: enough free CPU and memory for what the pod says it needs (its requests - the CPU and memory a pod reserves on a node), the pod's node rules match, the node's "keep out" marks (taints, 15.7) are allowed, ports are free.
  2. Score - rank the nodes that passed (spread copies of the same app apart, prefer less loaded nodes) and pick the highest.
  3. Bind - write a Binding: "this pod goes to that node". That sets spec.nodeName, and that is all the scheduler writes. It never starts anything.

When no node passes the filter, you get the scheduler's most useful message, as an Event (a log line attached to the pod, 15.11):

Warning  FailedScheduling  5s  default-scheduler  0/3 nodes are available: 1 node(s) had untolerated taint {node-role.kubernetes.io/control-plane: }, 2 Insufficient memory. preemption: 0/3 nodes are available: 1 Preemption is not helpful for scheduling, 2 No preemption victims found for incoming pod.

Read it as a tally per node: 0/3 nodes are available - none of 3 fit. 1 node(s) had untolerated taint {...control-plane} - cp-1 was excluded by its keep-out mark. 2 Insufficient memory - both workers lacked the memory the pod asked for. The preemption: part says evicting smaller pods would not help.

kube-controller-manager

A controller is a loop that watches one kind of object and keeps making reality match it (15.9 shows the loop in action). The kube-controller-manager is one program running dozens of them:

deployment controller    Deployment -> ReplicaSets
replicaset controller    ReplicaSet -> Pods
statefulset, daemonset, job, cronjob controllers
node controller          notices dead kubelets, marks nodes NotReady, evicts
endpoints / endpointslice controllers   Service -> the IPs of its ready pods
serviceaccount controller  a "default" ServiceAccount in every namespace
namespace controller     deletes everything inside a namespace being deleted
garbage collector        deletes objects whose owners are gone

Most names are kinds you meet later in this chapter (see the table in 15.3). The pattern is what matters: one controller per kind, each owning one job.

Its settings, extracted with jsonpath (a way to pick one field out of an object, like a jq path - 15.38 teaches it), then tr ',' '\n' puts each flag on its own line:

$ k get pod kube-controller-manager-cp-1 -n kube-system -o jsonpath='{.spec.containers[0].command}' | tr ',' '\n' | grep -E 'controllers|cluster-cidr|leader'
"--cluster-cidr=10.244.0.0/16"
"--controllers=*,bootstrapsigner,tokencleaner"
"--leader-elect=true"

--cluster-cidr is the pod network; --controllers=*,... = run all the default controllers plus two extras. --leader-elect=true matters with 3 control-plane nodes: all three run a controller-manager, but only one - the leader, the holder of a lock object called a Lease - acts. If it dies, another takes the lock. The scheduler works the same way.

cloud-controller-manager

When a cluster runs on a cloud provider, one more component talks to the cloud's own API: it creates a cloud load balancer (9.23) when you ask for a public Service, and deletes Node objects for VMs that no longer exist. On a managed cloud cluster the provider runs it for you. On this bare kubeadm cluster there is none - so a Service that asks for a cloud load balancer stays <pending> for ever here.

How the control plane runs itself: static pods

Chicken and egg: pods are created through the API server, but the API server is a pod. The answer is static pods: the kubelet (the node agent) on cp-1 reads YAML files from a directory, /etc/kubernetes/manifests/, and runs them directly - no API server involved:

/etc/kubernetes/manifests/
  etcd.yaml
  kube-apiserver.yaml
  kube-controller-manager.yaml
  kube-scheduler.yaml

The kubelet then creates a read-only copy in the API - a mirror pod - so you can see them. kube-apiserver-cp-1 is one. Recognise a mirror pod by its owner and a marker annotation (a note on the object, 15.26):

$ k get pod etcd-cp-1 -n kube-system -o jsonpath='{.metadata.ownerReferences[0].kind}/{.metadata.ownerReferences[0].name}{"\n"}'
Node/cp-1
$ k get pod etcd-cp-1 -n kube-system -o jsonpath='{.metadata.annotations.kubernetes\.io/config\.source}{"\n"}'
file

Deleting a mirror pod does nothing lasting - the kubelet recreates it from the file. To change the API server you edit the manifest file on cp-1 and the kubelet restarts it. Break the file and the API server disappears - the kubelet is only following the file, the same way systemd follows a unit file (2.1).

Where the rest of kube-system comes from

Not everything in kube-system is static. k get deploy,ds -n kube-system lists the Deployments and DaemonSets there (15.3's table):

$ k get deploy,ds -n kube-system
NAME                                      READY   UP-TO-DATE   AVAILABLE   AGE
deployment.apps/calico-kube-controllers   1/1     1            1           12d
deployment.apps/coredns                   2/2     2            2           12d
deployment.apps/metrics-server            1/1     1            1           12d

NAME                         DESIRED   CURRENT   READY   UP-TO-DATE   AVAILABLE   NODE SELECTOR            AGE
daemonset.apps/calico-node   3         3         3       3            3           kubernetes.io/os=linux   12d
daemonset.apps/kube-proxy    3         3         3       3            3           kubernetes.io/os=linux   12d

CoreDNS (the cluster DNS server) is an ordinary Deployment with 2 copies (2/2 ready). kube-proxy and calico-node are DaemonSets - one pod per node, including cp-1 (DESIRED 3). metrics-server collects CPU/memory usage for kubectl top. These are add-ons: they run on the cluster rather than making it.

What you can now do:

Why it helps

This is the model behind half of all Kubernetes troubleshooting. A pod stuck Pending with no events at all means the scheduler isn't running; with FailedScheduling, you read the per-node tally ("1 untolerated taint, 2 Insufficient memory"). An apply fails with "selector does not match template labels" or "unknown field spec.replica": that's the apiserver's validation, and you fix the manifest. "Why do we run 3 etcd members and not 4?" is a standard interview question with a one-table answer. In chapter 18 you'll break a static pod manifest and lose the apiserver, and knowing that kubectl can't help then, only crictl and the kubelet journal, is the difference between a five-minute and a five-hour fix.

FAQ

Why must etcd have an odd number of members?

Raft needs a majority, floor(n/2) + 1, to commit writes. 3 members have quorum 2 and tolerate 1 failure; 4 have quorum 3 and still tolerate only 1; 5 tolerate 2. An even member adds hardware that can fail without adding a failure you can survive. Production uses 3 or 5.

What happens if etcd loses quorum?

Running containers keep running, because kubelets and the runtime don't need etcd to keep processes alive. But nothing can change: no scheduling, scaling, deploying or deleting; every write to the apiserver fails. That's why etcd backups (etcdctl snapshot save) and quorum-aware maintenance are critical.

Does the scheduler start the pod?

No. It watches for pods without spec.nodeName, filters feasible nodes (requests versus allocatable, selectors, affinity, taints, volumes, ports), scores them, and writes a Binding that sets nodeName. That's all it writes. The kubelet on the chosen node sees a pod bound to it and starts it.

What is a static pod and why does the control plane use them?

A pod the kubelet runs directly from a YAML file in /etc/kubernetes/manifests, without the apiserver. It solves the chicken-and-egg problem: the apiserver itself is a pod. The kubelet creates a read-only mirror pod in the API so you can see it, owned by the Node, with config.source: file. Deleting the mirror does nothing lasting; edit the file to change it.

Why do control-plane pods show the node's IP?

They run with hostNetwork: true, using the node's network namespace. They have to: the apiserver, etcd, scheduler and controller-manager exist before the pod network (CNI) does, and other components reach them on the node's address, such as https://10.64.0.10:6443.

In an interview Junior

What does each control plane component do?

Four components, one job each, and one rule: only the API server reads and writes etcd.

On kubeadm they are static pods from /etc/kubernetes/manifests/ on cp-1; change them by editing those files.

Also asked: Why does a production etcd cluster have 3 or 5 members? · What is a static pod? · A pod stays Pending. Which component do you suspect first, and why?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.