Planned: cordon, drain, uncordon
The problem. Nodes go away in two ways: you take them away (patching, reboots) or they die. Planned, you want zero impact; unplanned, you need to know what Kubernetes does by itself, how long it waits, and why some pods never come back on their own.
What you need to know already: cordon, drain and PDBs (17.26), taints and tolerationSeconds (17.13), the kubelet on the node (18.1), StatefulSets (15.19).
$ kubectl cordon worker-1
node/worker-1 cordoned
$ kubectl get nodes
NAME STATUS ROLES AGE VERSION
cp-1 Ready control-plane 12d v1.34.1
worker-1 Ready,SchedulingDisabled <none> 12d v1.34.1
worker-2 Ready <none> 12d v1.34.1
cordon sets spec.unschedulable: true (and the taint node.kubernetes.io/unschedulable:NoSchedule). Nothing new lands there; nothing already there moves.
drain = cordon + evict every pod through the Eviction API, which respects PodDisruptionBudgets. It refuses three kinds of pod unless you tell it otherwise. On the lab's worker-1 it usually meets two of them: the DaemonSet pods, and metrics-server if it runs there (it keeps its scratch files in an emptyDir):
$ kubectl drain worker-1
node/worker-1 already cordoned
error: unable to drain node "worker-1" due to error: [cannot delete DaemonSet-managed Pods (use --ignore-daemonsets to ignore): kube-system/calico-node-2qs6n, kube-system/kube-proxy-v7zs8, cannot delete Pods with local storage (use --delete-emptydir-data to override): kube-system/metrics-server-4nwgl74xmz-2bsv2], continuing command...
There are pending nodes to be drained:
worker-1
...
| refusal | why | flag |
|---|---|---|
| DaemonSet pods | the DaemonSet would recreate them on the same node instantly | --ignore-daemonsets (leaves them running) |
| pods with emptyDir | its data is deleted with the pod | --delete-emptydir-data |
| pods with no controller | nothing will recreate them - evicting = deleting forever | --force |
| a PDB would be violated | budget says not now | none: drain waits and retries (fix the budget or add replicas) |
The standard maintenance command is therefore kubectl drain worker-1 --ignore-daemonsets --delete-emptydir-data, and --force is a decision, not a habit: find out what that bare pod is first. Mirror pods (static pods) are skipped - the apiserver cannot stop them anyway.
uncordon undoes the cordon. It does not move anything back: the pods that left stay where they went until something reschedules them (a rollout, a scale, another drain). A freshly uncordoned node is empty - expected. Give worker-1 back now (the drain above stopped before evicting anything):
$ kubectl uncordon worker-1
node/worker-1 uncordoned
Unplanned: the node stops answering
The kubelet posts status every 10s and renews its Lease (a tiny heartbeat object, one per node) in the kube-node-lease namespace. The node controller (in kube-controller-manager) watches those:
| time since last heartbeat | what you see |
|---|---|
| < 50s | Ready (the default --node-monitor-grace-period - how long the node controller waits for a heartbeat - is 50s since 1.32; it was 40s) |
| > 50s | Ready Unknown: Kubelet stopped posting node status. - STATUS NotReady |
| immediately then | taints node.kubernetes.io/unreachable:NoSchedule and :NoExecute |
| +300s | pods without a longer toleration are evicted (marked for deletion) |
Two flavours of NotReady, and they point to different places:
- Ready = False - the kubelet is alive and reporting a problem: runtime down, CNI missing (
container runtime network not ready), ... The taint isnode.kubernetes.io/not-ready. Go read what it says indescribe node. - Ready = Unknown - the kubelet is not reporting at all: kubelet stopped, machine down, network partition, apiserver unreachable from the node. The taint is
node.kubernetes.io/unreachable. Go to the machine.
# an illustration: worker-2 stopped answering a minute ago (the node incidents)
$ kubectl describe node worker-2 | grep -A6 Conditions
Conditions:
Type Status LastHeartbeatTime LastTransitionTime Reason Message
---- ------ ----------------- ------------------ ------ -------
MemoryPressure Unknown Tue, 22 Sep 2026 20:00:00 +0000 Tue, 22 Sep 2026 20:00:51 +0000 NodeStatusUnknown Kubelet stopped posting node status.
DiskPressure Unknown ...
Ready Unknown ...
Taint-based eviction: the 300 seconds
Every pod gets two tolerations added by the DefaultTolerationSeconds admission plugin:
# an illustration
kubectl get pod web-7c9 -o jsonpath='{.spec.tolerations}' | jq .
[
{ "effect": "NoExecute", "key": "node.kubernetes.io/not-ready", "operator": "Exists", "tolerationSeconds": 300 },
{ "effect": "NoExecute", "key": "node.kubernetes.io/unreachable", "operator": "Exists", "tolerationSeconds": 300 }
]
So pods on a dead node are left alone for 5 minutes (in case the node comes back), then evicted and recreated elsewhere by their controllers. Consequences:
- A node dying costs you up to ~50s + 300s of reduced capacity. For latency- sensitive services you can lower
tolerationSecondsin the pod spec (e.g. 30). - DaemonSet pods tolerate these taints with no limit - they are never evicted.
- StatefulSet pods are not replaced while the old pod object exists. The evicted pod goes
Terminatingand stays Terminating: the kubelet that should confirm the containers stopped is gone. A StatefulSet will not startdb-0elsewhere whiledb-0still exists - that is the at-most-one guarantee. Recovery: bring the node back, or, if it is truly dead,kubectl delete node(the pods go with it) or force-delete the pod once you are sure the machine is off.
# an illustration: six minutes after worker-2 died
$ kubectl get pods -o wide
NAME READY STATUS RESTARTS AGE NODE
web-w9m2pkd2r6-lng4v 1/1 Running 0 12s worker-1
web-w9m2pkd2r6-rlfg7 1/1 Terminating 0 6m19s worker-2
The Terminating one is the ghost on the dead node; the Running one is its replacement. When the node returns, its kubelet sees the deletion, kills the container and the ghost disappears.
Pods show Running on a NotReady node
Until eviction, pods on an unreachable node still show 1/1 Running: that is the last status the kubelet reported, not current truth. Nobody can know whether they run. The Pod's Ready condition goes False, so Services stop sending traffic to them - which is why a NotReady node does not usually mean errors for users, only less capacity.
Planned reboot: drain first
sudo reboot on a node without draining is an unplanned outage in slow motion: everything on it dies at once, and nothing is recreated elsewhere for 5+ minutes (it may come back first). Drain, reboot, wait for Ready, uncordon.
What you can now do
- Take a node out and back with cordon, drain (with the right flags) and uncordon.
- Read the timeline of a dead node: Unknown after ~50s, unreachable taints, eviction after 300s.
- Explain "Running" ghosts and Terminating pods on a dead node, and why StatefulSet pods wait.