What kubeadm is (and is not)
The problem. A new machine has to become a node, a broken one has to be rebuilt, an old one returned. Doing that means knowing what kubeadm writes, which kernel settings a node needs, and why a join command that worked yesterday fails today.
What you need to know already: the node's files and the kubelet unit (18.1), static pods (18.3), apt and packages (1.11), /etc/fstab and swap (the 2.36 incident), certificates and CAs (9.15), CertificateSigningRequests (17.43), drain (17.26).
kubeadm builds the control plane of a cluster and joins machines to it. It does not provision VMs, install a container runtime, install a CNI, or manage upgrades of your workloads. The division of labour:
| you | kubeadm |
|---|---|
| machines, OS, network between them | certificates (a CA + every component cert) |
| container runtime (containerd) configured for the systemd cgroup driver | kubeconfig files for admin, controller-manager, scheduler, kubelet |
| kubelet + kubeadm + kubectl packages | static pod manifests for etcd, apiserver, controller-manager, scheduler |
| the CNI (calico, cilium, flannel) after init | CoreDNS and kube-proxy, a bootstrap token, RBAC for joining |
That is exactly the shape of the lab: three VMs, containerd, the packages, then kubeadm init on cp-1, calico applied, kubeadm join on the workers.
Prerequisites on every node
kubeadm's preflight checks enforce most of these. Doing them by hand once is the best way to remember why each exists.
1. No swap (unless you configure the kubelet for it). The kubelet refuses to start with swap on by default (failSwapOn: true):
$ sudo swapoff -a # now
$ sudo sed -i '/ swap / s/^/#/' /etc/fstab # and after the next reboot
$ swapon --show # empty = off
The first line alone is the Block 1 incident: fixed until the next reboot.
2. IP forwarding. Pod traffic is routed through the node (the kernel must pass packets between interfaces, 8.11). sysctl is the tool that reads and sets kernel settings; files in /etc/sysctl.d/ are applied at every boot, sysctl --system applies them now:
$ echo 'net.ipv4.ip_forward = 1' | sudo tee /etc/sysctl.d/k8s.conf
$ sudo sysctl --system
* Applying /etc/sysctl.d/k8s.conf ...
net.ipv4.ip_forward = 1
sysctl -w net.ipv4.ip_forward=1 also works - until reboot. The file is what persists. Without it, kubeadm stops you:
error execution phase preflight: [preflight] Some fatal errors occurred:
[ERROR FileContent--proc-sys-net-ipv4-ip_forward]: /proc/sys/net/ipv4/ip_forward contents are not set to 1
3. Kernel modules (pieces of the kernel loaded on demand) overlay (the image filesystem, 10.3) and br_netfilter (so bridged pod traffic is seen by iptables, the kernel firewall, 16.3), loaded now with modprobe NAME and listed in /etc/modules-load.d/k8s.conf for the next boot (lsmod lists loaded modules):
# an illustration (no ▶): preparing and joining a new node (worker-3 is hypothetical)
$ printf 'overlay\nbr_netfilter\n' | sudo tee /etc/modules-load.d/k8s.conf
sudo modprobe overlay && sudo modprobe br_netfilter
lsmod | grep br_netfilter
br_netfilter 32768 0
4. A container runtime, running and enabled. Here containerd 2.1. The one setting that bites everybody: containerd and the kubelet must use the same cgroup driver (the program that creates and manages the cgroups - systemd, or the runtime itself). kubeadm configures the kubelet for systemd (cgroupDriver: systemd in /var/lib/kubelet/config.yaml), so containerd needs SystemdCgroup = true in /etc/containerd/config.toml. A mismatch gives a node that joins, then pods and even the static pods restart over and over. (simulator) The lab's containerd is already configured this way; the mismatch itself is not simulated.
5. The packages, from the right repository. Since 2023 Kubernetes packages live at pkgs.k8s.io, with one repository per minor version:
# an illustration (no ▶): preparing and joining a new node (worker-3 is hypothetical)
cat /etc/apt/sources.list.d/kubernetes.list
deb [signed-by=/etc/apt/keyrings/kubernetes-apt-keyring.gpg] https://pkgs.k8s.io/core:/stable:/v1.34/deb/ /
A node pointed at /v1.34/deb/ can only ever install 1.34.x. That is deliberate: apt upgrade cannot silently jump a minor version. It also means an upgrade to 1.35 starts with editing this line. The old apt.kubernetes.io repository is frozen since September 2023 - old blog posts that use it will not work.
(simulator) The signing key is already in /etc/apt/keyrings/ on the lab VMs. On a real machine: curl -fsSL https://pkgs.k8s.io/core:/stable:/v1.34/deb/Release.key | sudo gpg --dearmor -o /etc/apt/keyrings/kubernetes-apt-keyring.gpg.
apt-cache madison PKG lists every version of a package the repositories offer; apt-get install PKG=VERSION installs an exact one; apt-mark hold freezes a package so upgrades skip it:
# an illustration (no ▶): preparing and joining a new node (worker-3 is hypothetical)
$ sudo apt-get update
...
Get:5 https://prod-cdn.packages.k8s.io/repositories/isv:/kubernetes:/core:/stable:/v1.34/deb InRelease [1224 B]
apt-cache madison kubeadm
kubeadm | 1.34.4-1.1 | https://pkgs.k8s.io/core:/stable:/v1.34/deb Packages
kubeadm | 1.34.2-1.1 | https://pkgs.k8s.io/core:/stable:/v1.34/deb Packages
kubeadm | 1.34.1-1.1 | https://pkgs.k8s.io/core:/stable:/v1.34/deb Packages
kubeadm | 1.34.0-1.1 | https://pkgs.k8s.io/core:/stable:/v1.34/deb Packages
sudo apt-get install -y kubelet=1.34.1-1.1 kubeadm=1.34.1-1.1 kubectl=1.34.1-1.1
sudo apt-mark hold kubelet kubeadm kubectl
kubelet set on hold.
kubeadm set on hold.
kubectl set on hold.
Pin the version to match the cluster (=1.34.1-1.1), and hold the three packages so a routine apt upgrade never upgrades Kubernetes behind your back. (The patch versions listed here are what this lab's mirror carries.)
Right after install the kubelet is enabled and crash-loops every 10 seconds:
# on a node right after the install, before kubeadm init or join
$ systemctl status kubelet
Active: activating (auto-restart) (Result: exit-code) since ...; 4s ago
$ journalctl -u kubelet -n 3 --no-pager
... "command failed" err="failed to load kubelet config file, path: /var/lib/kubelet/config.yaml, error: ... no such file or directory"
That is expected: the kubelet has no config until kubeadm init or kubeadm join writes one. The kubeadm docs say it in those words - the kubelet "is restarting every few seconds, as it waits in a crashloop for kubeadm to tell it what to do".
kubeadm init, phase by phase
On cp-1 the cluster was created with sudo kubeadm init --pod-network-cidr=10.244.0.0/16. The output is a tour of everything this chapter covers; read it once properly:
[init] Using Kubernetes version: v1.34.1
[preflight] Running pre-flight checks
[preflight] Pulling images required for setting up a Kubernetes cluster
[certs] Using certificateDir folder "/etc/kubernetes/pki"
[certs] Generating "ca" certificate and key
[certs] Generating "apiserver" certificate and key
[certs] apiserver serving cert is signed for DNS names [cp-1 kubernetes kubernetes.default kubernetes.default.svc kubernetes.default.svc.cluster.local] and IPs [10.96.0.1 10.64.0.10]
[certs] Generating "apiserver-kubelet-client" certificate and key
[certs] Generating "front-proxy-ca" certificate and key
[certs] Generating "front-proxy-client" certificate and key
[certs] Generating "etcd/ca" certificate and key
[certs] Generating "etcd/server" certificate and key
[certs] Generating "etcd/peer" certificate and key
[certs] Generating "etcd/healthcheck-client" certificate and key
[certs] Generating "apiserver-etcd-client" certificate and key
[certs] Generating "sa" key and public key
[kubeconfig] Using kubeconfig folder "/etc/kubernetes"
[kubeconfig] Writing "admin.conf" kubeconfig file
[kubeconfig] Writing "super-admin.conf" kubeconfig file
[kubeconfig] Writing "kubelet.conf" kubeconfig file
[kubeconfig] Writing "controller-manager.conf" kubeconfig file
[kubeconfig] Writing "scheduler.conf" kubeconfig file
[etcd] Creating static Pod manifest for local etcd in "/etc/kubernetes/manifests"
[control-plane] Using manifest folder "/etc/kubernetes/manifests"
[control-plane] Creating static Pod manifest for "kube-apiserver"
[control-plane] Creating static Pod manifest for "kube-controller-manager"
[control-plane] Creating static Pod manifest for "kube-scheduler"
[kubelet-start] Writing kubelet environment file with flags to file "/var/lib/kubelet/kubeadm-flags.env"
[kubelet-start] Writing kubelet configuration to file "/var/lib/kubelet/config.yaml"
[kubelet-start] Starting the kubelet
[wait-control-plane] Waiting for the kubelet to boot up the control plane as static Pods from directory "/etc/kubernetes/manifests"
[upload-config] Storing the configuration used in ConfigMap "kubeadm-config" in the "kube-system" Namespace
[kubelet] Creating a ConfigMap "kubelet-config" in namespace kube-system with the configuration for the kubelets in the cluster
[mark-control-plane] Marking the node cp-1 as control-plane by adding the labels: [node-role.kubernetes.io/control-plane node.kubernetes.io/exclude-from-external-load-balancers]
[mark-control-plane] Marking the node cp-1 as control-plane by adding the taints [node-role.kubernetes.io/control-plane:NoSchedule]
[bootstrap-token] Using token: abcdef.0123456789abcdef
[bootstrap-token] Configuring bootstrap tokens, cluster-info ConfigMap, RBAC Roles
[kubelet-finalize] Updating "/etc/kubernetes/kubelet.conf" to point to a rotatable kubelet client certificate and key
[addons] Applied essential addon: CoreDNS
[addons] Applied essential addon: kube-proxy
Your Kubernetes control-plane has initialized successfully!
To start using your cluster, you need to run the following as a regular user:
mkdir -p $HOME/.kube
sudo cp -i /etc/kubernetes/admin.conf $HOME/.kube/config
sudo chown $(id -u):$(id -g) $HOME/.kube/config
You should now deploy a pod network to the cluster.
Run "kubectl apply -f [podnetwork].yaml" with one of the options listed at:
https://kubernetes.io/docs/concepts/cluster-administration/addons/
Then you can join any number of worker nodes by running the following on each as root:
kubeadm join 10.64.0.10:6443 --token abcdef.0123456789abcdef \
--discovery-token-ca-cert-hash sha256:9fe8ef7e62af965c...
(The exact wait/health-check lines change between minor versions; the phases do not.) Map it onto what you already know:
- [certs] - the whole PKI (public key infrastructure: the CAs plus all the certificates they signed) in
/etc/kubernetes/pki. Three CAs (ca,etcd/ca,front-proxy-ca), leaf certs (the ones actually used by a component, not a CA) valid one year. The apiserver cert's SANs (Subject Alternative Names, 9.17) are every name and IP clients use - if you later add a load balancer name, the cert must be re-issued with it. - [kubeconfig] - kubeconfigs that embed client certificates.
admin.confiskubeadm:cluster-admins;super-admin.confissystem:masters- the break-glass one that bypasses RBAC, keep it off workstations. - [etcd]/[control-plane] - four files in
/etc/kubernetes/manifests; the kubelet does the rest (previous lesson). - [mark-control-plane] - the taint that keeps your pods off cp-1.
- [bootstrap-token] - a bootstrap token (a short-lived password for joining nodes), valid 24 hours.
- addons - CoreDNS and kube-proxy. Not the CNI: until you apply one, the nodes stay NotReady ("cni plugin not initialized") and CoreDNS stays Pending.
kubeadm join
kubeadm join 10.64.0.10:6443 --token 7x3kqp.u2m9d0e1f4g5h6j7 --discovery-token-ca-cert-hash sha256:9fe8ef7e...
Two secrets, two directions of trust:
--tokenauthenticates the node to the cluster. It is a bootstrap token (a Secretbootstrap-token-<id>in kube-system) with a TTL.--discovery-token-ca-cert-hashauthenticates the cluster to the node: the node downloads the CA from the publiccluster-infoConfigMap and checks its public key hash against this pin (a pin = a hash you trust in advance), so a man-in-the-middle (someone between the node and the cluster) cannot hand it a fake CA. Leaving it out is possible (--discovery-token-unsafe-skip-ca-verification) and exactly as unsafe as it sounds.
Then the kubelet does a TLS bootstrap (getting its own certificate automatically): it uses the token to send a CertificateSigningRequest, the controller-manager auto-approves it, the kubelet gets its own client certificate (/var/lib/kubelet/pki/kubelet-client-current.pem) and writes /etc/kubernetes/kubelet.conf. From then on the token is not needed.
This node has joined the cluster:
* Certificate signing request was sent to apiserver and a response was received.
* The Kubelet was informed of the new secure connection details.
Run 'kubectl get nodes' on the control-plane to see this node join the cluster.
The node appears NotReady for a few seconds - until the calico-node pod of the DaemonSet starts there and writes the CNI config - then Ready.
The join command from yesterday
The init token expires after 24h. The day after, the command in your notes fails like this:
error execution phase preflight: couldn't validate the identity of the API Server: could not find a JWS signature in the cluster-info ConfigMap for token ID "abcdef"
Nothing is wrong with the node. Make a fresh one on cp-1:
# an illustration (no ▶): preparing and joining a new node (worker-3 is hypothetical)
sudo kubeadm token create --print-join-command
kubeadm join 10.64.0.10:6443 --token a9e215.3c911a3b9c0b5ec1 --discovery-token-ca-cert-hash sha256:9fe8ef7e62af965c...
sudo kubeadm token list
TOKEN TTL EXPIRES USAGES DESCRIPTION EXTRA GROUPS
a9e215.3c911a3b9c0b5ec1 23h 2026-09-23T20:00:03Z authentication,signing <none> system:bootstrappers:kubeadm:default-node-token
A handy pattern from oncall-lab, since ssh passes output back:
# an illustration (no ▶): preparing and joining a new node (worker-3 is hypothetical)
$ JOIN=$(ssh cp-1 sudo kubeadm token create --print-join-command)
ssh worker-3 "sudo $JOIN"
A failed join leaves debris
If a join fails half way (the kubelet never became healthy, say because swap was on), files are already written. The next attempt stops in preflight:
[ERROR FileAvailable--etc-kubernetes-pki-ca.crt]: /etc/kubernetes/pki/ca.crt already exists
The fix is sudo kubeadm reset -f on that node, fix the cause, join again. And reset has limits - it prints them:
The reset process does not clean CNI configuration. To do so, you must remove /etc/cni/net.d
The reset process does not reset or clean up iptables rules or IPVS tables.
The reset process does not clean your kubeconfig files and you must remove them manually.
Removing a node properly
The order matters - workloads first, then the object, then the machine:
# an illustration (no ▶): preparing and joining a new node (worker-3 is hypothetical)
kubectl drain worker-3 --ignore-daemonsets --delete-emptydir-data # move the pods away
kubectl delete node worker-3 # forget it in the apiserver
ssh worker-3 sudo kubeadm reset -f # undo join on the machine
Delete the node object without resetting and the still-running kubelet re-registers it within seconds. Reset without draining and the pods just die with the node.
What you can now do
- Prepare a node persistently: swap, ip_forward, modules, runtime, pinned and held packages.
- Read
kubeadm initoutput phase by phase and explain the join's two secrets. - Recover a failed or stale join (
kubeadm reset,token create --print-join-command) and remove a node in the right order.