OnCallReady

Lesson 15.7 · Kubernetes: Architecture & Workloads · 17 min read

The node: kubelet, CRI, CNI, kube-proxy, CoreDNS

In plain words

Each kitchen station in a restaurant has a station chef who reads the orders assigned to that station and cooks them, a stove that actually does the heating, a plumber who connects each new sink to the water, and a waiter who knows which table each dish goes to. The station chef tells the head office every few seconds "I'm here and working".

On each Kubernetes node, the kubelet is the station chef: a systemd service that watches for pods bound to its node, asks containerd (the stove) over CRI to run them, lets the CNI plugin (Calico, the plumber) give each pod an IP, runs probes, and posts node status as a heartbeat. kube-proxy programs iptables so Service IPs reach pods. kubectl describe node shows Conditions, Capacity versus Allocatable, and allocated requests.

Why know the node

The control plane only decides. Everything that actually runs - every container, every pod IP - is made by a few programs on each node. When a node goes NotReady, or pods are stuck in ContainerCreating, those programs are where you look, and you already know how to debug most of them: one of them is just a systemd service.

What you need to know already: the control plane (15.5), containers, containerd and runc (10.3), network namespaces and veth pairs (11.15), systemd status and journalctl (2.9, 2.30), the node-NotReady incident (2.36), DNS and /etc/resolv.conf (8.16), CIDR (8.3).

What runs on every node

kubelet        systemd service on the host. Watches the apiserver for pods bound
               to THIS node, makes them happen, reports status back.
containerd     the container runtime. The kubelet talks to it over CRI
               (Container Runtime Interface, gRPC on a unix socket).
CNI plugin     called by the runtime when a pod sandbox is created: gives the pod
               a network namespace, an interface and an IP. Here: Calico.
kube-proxy     DaemonSet. Watches Services/EndpointSlices and programs iptables
               (or IPVS/nftables) so a Service IP DNATs to a pod IP.

In plain words:

CoreDNS is not a node component - it is a normal Deployment somewhere in the cluster - but every pod's /etc/resolv.conf (8.16) points at it, so every name lookup a pod does goes through it.

The kubelet

For each pod the scheduler bound to its node, the kubelet:

  1. asks the runtime (via CRI) for a pod sandbox - the pod's own network namespace, held open by a tiny "pause" container; the CNI plugin gives it an IP here
  2. mounts the pod's volumes (files the pod needs: settings, passwords, disks)
  3. pulls the images (like docker pull)
  4. runs the pod's setup containers (init containers, 15.35) in order, each to completion
  5. creates and starts the app containers
  6. runs the health checks (probes) and restarts containers that exit, waiting longer each time (backoff, 15.14) - like systemd's Restart= and RestartSec= (2.10)
  7. writes pod status and node status (a regular "I'm alive" report, the heartbeat) back to the API server

It also runs the static pods from /etc/kubernetes/manifests (15.5).

And it is not a pod: it is kubelet.service on the host. When it dies the node goes NotReady, and you debug it with systemctl status kubelet and journalctl -u kubelet - exactly what you did in the node-NotReady incident (2.36).

Later (Ch 17): probes get a full lesson - how the kubelet decides a container is healthy or ready.

CRI, and why "Docker" is not in the picture

Kubernetes stopped talking to Docker in v1.24. The kubelet speaks CRI to containerd directly (or to CRI-O, another runtime). Images built with docker build still run - they are standard OCI images (10.53) - but on a node, docker ps shows nothing, because Docker is not installed.

The node-side tool is crictl, a small CLI that talks CRI to the runtime:

sudo crictl ps          # running containers, from containerd via CRI
sudo crictl pods        # pod sandboxes
sudo crictl logs <id>   # works even when the apiserver is down

It feels like the docker CLI; you run it with sudo on the node, not on your workstation.

CNI

When the sandbox is created, containerd runs the CNI plugin program from /opt/cni/bin, with its settings from /etc/cni/net.d. The plugin:

Each node gets its own /24 slice of 10.244.0.0/16. This jsonpath (15.38 teaches the syntax: range loops over the nodes, \t and \n add tabs and newlines) prints name and slice per node:

$ k get nodes -o jsonpath='{range .items[*]}{.metadata.name}{"\t"}{.spec.podCIDR}{"\n"}{end}'
cp-1	10.244.0.0/24
worker-1	10.244.1.0/24
worker-2	10.244.2.0/24

(Calico actually hands out addresses in smaller /26 blocks from its own pool; the node's podCIDR is what the controller-manager assigned. Same idea.)

If the CNI plugin is missing or broken, pods never get past ContainerCreating, and the node itself reports NotReady. The kubelet records an Event (a log line on the pod, 15.11) like:

Warning  FailedCreatePodSandBox  3s  kubelet  Failed to create pod sandbox: rpc error: code = Unknown desc = failed to setup network for sandbox "4f1c...": plugin type="calico" failed (add): stat /var/lib/calico/nodename: no such file or directory: check that the calico/node container is running and has mounted /var/lib/calico/

Read it from the right: the Calico plugin could not find a file its own node pod writes - so the Calico pod on that node is not running.

describe node, decoded

kubectl describe prints everything about one object in a readable layout (get -o yaml is the raw version). On a node:

$ k describe node worker-1
Name:               worker-1
Roles:              <none>
Labels:             beta.kubernetes.io/arch=arm64
                    kubernetes.io/arch=arm64
                    kubernetes.io/hostname=worker-1
                    kubernetes.io/os=linux
Annotations:        kubeadm.alpha.kubernetes.io/cri-socket: unix:///var/run/containerd/containerd.sock
                    projectcalico.org/IPv4Address: 10.64.0.11/24
...
Taints:             <none>
Unschedulable:      false
...
Conditions:
  Type             Status  LastHeartbeatTime                 LastTransitionTime                Reason                       Message
  ----             ------  -----------------                 ------------------                ------                       -------
  MemoryPressure   False   Wed, 23 Sep 2026 10:20:31 +0000   Thu, 10 Sep 2026 16:47:03 +0000   KubeletHasSufficientMemory   kubelet has sufficient memory available
  DiskPressure     False   ...                                                                 KubeletHasNoDiskPressure     kubelet has no disk pressure
  PIDPressure      False   ...                                                                 KubeletHasSufficientPID      kubelet has sufficient PID available
  Ready            True    ...                                                                 KubeletReady                 kubelet is posting ready status
Addresses:
  InternalIP: 10.64.0.11
  Hostname:   worker-1
Capacity:
  cpu:                2
  memory:             4005528Ki
  pods:               110
Allocatable:
  cpu:                2
  memory:             3903128Ki
  pods:               110
...
Non-terminated Pods:          (4 in total)
  Namespace     Name                              CPU Requests   CPU Limits   Memory Requests   Memory Limits   Age
  ---------     ----                              ------------   ----------   ---------------   -------------   ---
  kube-system   calico-node-2qs6n                 250m (12%)     0 (0%)       0 (0%)            0 (0%)          12d
  kube-system   coredns-d6qs9mxzfp-7n9b6          100m (5%)      0 (0%)       70Mi (1%)         170Mi (4%)      12d
...
Allocated resources:
  Resource            Requests     Limits
  --------            --------     ------
  cpu                 450m (22%)   0m (0%)
  memory              270Mi (7%)   170Mi (4%)

Top part: Labels are name tags (15.26) - here the CPU architecture, OS and hostname. Annotations are free-form notes - here the CRI socket path. Taints are keep-out marks (next section); Unschedulable: false means the node accepts new pods.

The four sections that matter:

Requests and limits are set on each container; this chapter only reads them.

kubectl top shows actual usage, measured by metrics-server (15.5):

$ k top nodes
NAME       CPU(cores)   CPU(%)   MEMORY(bytes)   MEMORY(%)
cp-1       176m         9%       1376Mi          36%
worker-1   64m          3%       662Mi           17%
worker-2   60m          3%       641Mi           17%

Taints on the control plane

A taint is a keep-out mark on a node: pods are not scheduled there unless they carry a matching toleration ("I accept that mark"). grep Taints keeps just that line of the describe output:

$ k describe node cp-1 | grep Taints
Taints:             node-role.kubernetes.io/control-plane:NoSchedule

Read it as key:effect. The key says "this is a control-plane node"; the effect NoSchedule = do not put new pods here. kubeadm adds it so your apps stay off the brain. A pod lands on cp-1 only if it tolerates the taint - kube-proxy, calico-node and coredns do. That is why your Deployments spread over two workers, not three.

What you can now do:

Why it helps

When a node goes NotReady at 3am, this lesson tells you where to look: the kubelet is not a pod, so systemctl status kubelet and journalctl -u kubelet on the node, exactly your chapter 1 skills. Pods stuck in ContainerCreating with FailedCreatePodSandBox point at the CNI. A developer asks why their pod doesn't fit on a node that kubectl top shows at 3% CPU; the answer is that the scheduler uses requests against Allocatable, not usage. And docker ps on a node showing nothing surprises people until they know the kubelet speaks CRI to containerd, and the tool is crictl.

FAQ

Why does docker ps show nothing on a Kubernetes node?

Kubernetes removed dockershim in 1.24; the kubelet talks to containerd (or CRI-O) directly over CRI. Docker, if installed, is a separate runtime that knows nothing about Kubernetes pods. Use sudo crictl ps, crictl pods and crictl logs on the node; they work even when the apiserver is down.

What is the difference between Capacity and Allocatable?

Capacity is the whole machine: CPU, memory and pod slots. Allocatable is what's left for pods after kube-reserved, system-reserved and the eviction threshold, such as the default 100Mi memory hard-eviction margin. The scheduler places pods against Allocatable minus the requests of pods already there.

Why doesn't my pod fit when the node is almost idle?

Scheduling uses resource requests, not actual usage. "Allocated resources" in kubectl describe node sums the requests of pods on the node; if a new pod's requests don't fit in what's left of Allocatable, it won't schedule there, even at 3% real CPU. kubectl top node shows usage from metrics-server, which is a different number.

What happens when a kubelet stops reporting?

The node's conditions go Unknown with "Kubelet stopped posting node status" after the grace period, and the node is marked NotReady. After a further timeout, pods on it get evicted via the unreachable taint and controllers recreate them elsewhere (StatefulSet pods behave more carefully). Debug on the node with systemctl status kubelet and journalctl -u kubelet.

Why do my Deployments only use two of three nodes?

kubeadm taints the control plane with node-role.kubernetes.io/control-plane:NoSchedule, so only pods that tolerate it land there, like kube-proxy, calico-node and CoreDNS. Your workloads spread over the two workers. You'll see the same in DESIRED 2 for a DaemonSet without that toleration.

In an interview Junior

What does the kubelet do?

The kubelet is the node agent - a systemd service on every node, not a pod - and the only component that actually starts containers. It watches the API server for pods bound to its node, and for each one:

  1. asks the runtime through CRI (containerd) for a pod sandbox; the CNI plugin gives it an IP
  2. mounts the volumes, pulls the images
  3. runs the init containers in order, then starts the app containers
  4. runs the probes and restarts exited containers with backoff
  5. writes pod status and the node's heartbeat back to the API server

It also runs the static pods from /etc/kubernetes/manifests.

When it stops, the node goes NotReady (describe node conditions go Unknown: "Kubelet stopped posting node status"). Debug it like any service: systemctl status kubelet, journalctl -u kubelet. On the node, sudo crictl ps lists containers - there is no docker.

Also asked: A node shows NotReady. How do you troubleshoot it? · What are CRI, CNI and kube-proxy? · What is the difference between allocatable and capacity on a node?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.