Why know the node
The control plane only decides. Everything that actually runs - every container, every pod IP - is made by a few programs on each node. When a node goes NotReady, or pods are stuck in ContainerCreating, those programs are where you look, and you already know how to debug most of them: one of them is just a systemd service.
What you need to know already: the control plane (15.5), containers, containerd and runc (10.3), network namespaces and veth pairs (11.15), systemd status and journalctl (2.9, 2.30), the node-NotReady incident (2.36), DNS and /etc/resolv.conf (8.16), CIDR (8.3).
What runs on every node
kubelet systemd service on the host. Watches the apiserver for pods bound
to THIS node, makes them happen, reports status back.
containerd the container runtime. The kubelet talks to it over CRI
(Container Runtime Interface, gRPC on a unix socket).
CNI plugin called by the runtime when a pod sandbox is created: gives the pod
a network namespace, an interface and an IP. Here: Calico.
kube-proxy DaemonSet. Watches Services/EndpointSlices and programs iptables
(or IPVS/nftables) so a Service IP DNATs to a pod IP.
In plain words:
- kubelet - the node agent. The only component that actually starts containers.
- containerd - the container runtime (10.3), the same one under Docker.
- CRI (Container Runtime Interface) - the standard API the kubelet uses to tell a runtime "start this, stop that". gRPC is a kind of API call; the unix socket is a file-shaped connection on the same machine, like Docker's
/var/run/docker.sock(10.1). - CNI (Container Network Interface) - the standard for network plugins. The plugin (here Calico) gives each pod its own network and IP.
- kube-proxy - makes Service addresses work on this node by writing iptables rules (the kernel's packet-rewriting table): traffic sent to a Service IP gets its destination rewritten (DNAT) to one real pod IP. Services are Ch 16; for now: kube-proxy is the per-node piece that makes them work.
CoreDNS is not a node component - it is a normal Deployment somewhere in the cluster - but every pod's /etc/resolv.conf (8.16) points at it, so every name lookup a pod does goes through it.
The kubelet
For each pod the scheduler bound to its node, the kubelet:
- asks the runtime (via CRI) for a pod sandbox - the pod's own network namespace, held open by a tiny "pause" container; the CNI plugin gives it an IP here
- mounts the pod's volumes (files the pod needs: settings, passwords, disks)
- pulls the images (like
docker pull) - runs the pod's setup containers (init containers, 15.35) in order, each to completion
- creates and starts the app containers
- runs the health checks (probes) and restarts containers that exit, waiting longer each time (backoff, 15.14) - like systemd's
Restart=andRestartSec=(2.10) - writes pod status and node status (a regular "I'm alive" report, the heartbeat) back to the API server
It also runs the static pods from /etc/kubernetes/manifests (15.5).
And it is not a pod: it is kubelet.service on the host. When it dies the node goes NotReady, and you debug it with systemctl status kubelet and journalctl -u kubelet - exactly what you did in the node-NotReady incident (2.36).
Later (Ch 17): probes get a full lesson - how the kubelet decides a container is healthy or ready.
CRI, and why "Docker" is not in the picture
Kubernetes stopped talking to Docker in v1.24. The kubelet speaks CRI to containerd directly (or to CRI-O, another runtime). Images built with docker build still run - they are standard OCI images (10.53) - but on a node, docker ps shows nothing, because Docker is not installed.
The node-side tool is crictl, a small CLI that talks CRI to the runtime:
sudo crictl ps # running containers, from containerd via CRI
sudo crictl pods # pod sandboxes
sudo crictl logs <id> # works even when the apiserver is down
It feels like the docker CLI; you run it with sudo on the node, not on your workstation.
CNI
When the sandbox is created, containerd runs the CNI plugin program from /opt/cni/bin, with its settings from /etc/cni/net.d. The plugin:
- creates a veth pair (a virtual cable, 11.15) and puts one end in the pod
- gives it an address from this node's slice of the pod network
- sets up routes so pods on other nodes are reachable
Each node gets its own /24 slice of 10.244.0.0/16. This jsonpath (15.38 teaches the syntax: range loops over the nodes, \t and \n add tabs and newlines) prints name and slice per node:
$ k get nodes -o jsonpath='{range .items[*]}{.metadata.name}{"\t"}{.spec.podCIDR}{"\n"}{end}'
cp-1 10.244.0.0/24
worker-1 10.244.1.0/24
worker-2 10.244.2.0/24
(Calico actually hands out addresses in smaller /26 blocks from its own pool; the node's podCIDR is what the controller-manager assigned. Same idea.)
If the CNI plugin is missing or broken, pods never get past ContainerCreating, and the node itself reports NotReady. The kubelet records an Event (a log line on the pod, 15.11) like:
Warning FailedCreatePodSandBox 3s kubelet Failed to create pod sandbox: rpc error: code = Unknown desc = failed to setup network for sandbox "4f1c...": plugin type="calico" failed (add): stat /var/lib/calico/nodename: no such file or directory: check that the calico/node container is running and has mounted /var/lib/calico/
Read it from the right: the Calico plugin could not find a file its own node pod writes - so the Calico pod on that node is not running.
describe node, decoded
kubectl describe prints everything about one object in a readable layout (get -o yaml is the raw version). On a node:
$ k describe node worker-1
Name: worker-1
Roles: <none>
Labels: beta.kubernetes.io/arch=arm64
kubernetes.io/arch=arm64
kubernetes.io/hostname=worker-1
kubernetes.io/os=linux
Annotations: kubeadm.alpha.kubernetes.io/cri-socket: unix:///var/run/containerd/containerd.sock
projectcalico.org/IPv4Address: 10.64.0.11/24
...
Taints: <none>
Unschedulable: false
...
Conditions:
Type Status LastHeartbeatTime LastTransitionTime Reason Message
---- ------ ----------------- ------------------ ------ -------
MemoryPressure False Wed, 23 Sep 2026 10:20:31 +0000 Thu, 10 Sep 2026 16:47:03 +0000 KubeletHasSufficientMemory kubelet has sufficient memory available
DiskPressure False ... KubeletHasNoDiskPressure kubelet has no disk pressure
PIDPressure False ... KubeletHasSufficientPID kubelet has sufficient PID available
Ready True ... KubeletReady kubelet is posting ready status
Addresses:
InternalIP: 10.64.0.11
Hostname: worker-1
Capacity:
cpu: 2
memory: 4005528Ki
pods: 110
Allocatable:
cpu: 2
memory: 3903128Ki
pods: 110
...
Non-terminated Pods: (4 in total)
Namespace Name CPU Requests CPU Limits Memory Requests Memory Limits Age
--------- ---- ------------ ---------- --------------- ------------- ---
kube-system calico-node-2qs6n 250m (12%) 0 (0%) 0 (0%) 0 (0%) 12d
kube-system coredns-d6qs9mxzfp-7n9b6 100m (5%) 0 (0%) 70Mi (1%) 170Mi (4%) 12d
...
Allocated resources:
Resource Requests Limits
-------- -------- ------
cpu 450m (22%) 0m (0%)
memory 270Mi (7%) 170Mi (4%)
Top part: Labels are name tags (15.26) - here the CPU architecture, OS and hostname. Annotations are free-form notes - here the CRI socket path. Taints are keep-out marks (next section); Unschedulable: false means the node accepts new pods.
The four sections that matter:
- Conditions - the kubelet's self-report.
Ready=Trueplus three "pressure" checks (low memory, low disk, too many processes - allFalseis good).LastHeartbeatTimeis the last report. If the kubelet stops reporting, all four goUnknownwith "Kubelet stopped posting node status." - Capacity vs Allocatable - capacity is the whole machine; allocatable is what is left for pods after the node keeps some back for itself. Here about 100Mi of memory is held back. The scheduler uses allocatable.
- Non-terminated Pods - who runs here, and what each requested. Requests are the CPU/memory a pod reserves (
250m= 250 millicores = a quarter of one CPU;70Mi= 70 mebibytes). Limits are the ceiling - a cgroup limit likeMemoryMax=(5.11). - Allocated resources - the sum of requests vs allocatable. This, not actual usage, decides whether a new pod fits. A node can be 22% "allocated" and 90% actually busy, or the reverse.
Requests and limits are set on each container; this chapter only reads them.
kubectl top shows actual usage, measured by metrics-server (15.5):
$ k top nodes
NAME CPU(cores) CPU(%) MEMORY(bytes) MEMORY(%)
cp-1 176m 9% 1376Mi 36%
worker-1 64m 3% 662Mi 17%
worker-2 60m 3% 641Mi 17%
Taints on the control plane
A taint is a keep-out mark on a node: pods are not scheduled there unless they carry a matching toleration ("I accept that mark"). grep Taints keeps just that line of the describe output:
$ k describe node cp-1 | grep Taints
Taints: node-role.kubernetes.io/control-plane:NoSchedule
Read it as key:effect. The key says "this is a control-plane node"; the effect NoSchedule = do not put new pods here. kubeadm adds it so your apps stay off the brain. A pod lands on cp-1 only if it tolerates the taint - kube-proxy, calico-node and coredns do. That is why your Deployments spread over two workers, not three.
What you can now do:
- name what runs on every node and what each piece does
- debug the kubelet as the systemd service it is
- read
describe node: conditions, allocatable, requests, taints