OnCallReady

Lesson 18.25 · Kubernetes: Cluster Operations & Troubleshooting · 13 min read

Node lifecycle: cordon, drain, NotReady and taint-based eviction

In plain words

Imagine a hotel floor that needs repairs. First you put a "no new guests" sign on it (cordon). Then you kindly move each guest to another floor, but never so many from one family at once that the family can't cope (drain, respecting PDBs). After the repairs, you take the sign down (uncordon), but the guests don't move back by themselves.

Now imagine a floor where the phone line suddenly goes dead. The front desk waits a bit (about 50 seconds), marks the floor "unknown", then waits five more minutes in case the line comes back, and only then books the guests into other rooms. That's the difference between planned node maintenance (cordon, drain, uncordon) and an unplanned failure (NotReady, taints, eviction after tolerationSeconds: 300).

Planned: cordon, drain, uncordon

The problem. Nodes go away in two ways: you take them away (patching, reboots) or they die. Planned, you want zero impact; unplanned, you need to know what Kubernetes does by itself, how long it waits, and why some pods never come back on their own.

What you need to know already: cordon, drain and PDBs (17.26), taints and tolerationSeconds (17.13), the kubelet on the node (18.1), StatefulSets (15.19).

$ kubectl cordon worker-1
node/worker-1 cordoned
$ kubectl get nodes
NAME       STATUS                     ROLES           AGE   VERSION
cp-1       Ready                      control-plane   12d   v1.34.1
worker-1   Ready,SchedulingDisabled   <none>          12d   v1.34.1
worker-2   Ready                      <none>          12d   v1.34.1

cordon sets spec.unschedulable: true (and the taint node.kubernetes.io/unschedulable:NoSchedule). Nothing new lands there; nothing already there moves.

drain = cordon + evict every pod through the Eviction API, which respects PodDisruptionBudgets. It refuses three kinds of pod unless you tell it otherwise. On the lab's worker-1 it usually meets two of them: the DaemonSet pods, and metrics-server if it runs there (it keeps its scratch files in an emptyDir):

$ kubectl drain worker-1
node/worker-1 already cordoned
error: unable to drain node "worker-1" due to error: [cannot delete DaemonSet-managed Pods (use --ignore-daemonsets to ignore): kube-system/calico-node-2qs6n, kube-system/kube-proxy-v7zs8, cannot delete Pods with local storage (use --delete-emptydir-data to override): kube-system/metrics-server-4nwgl74xmz-2bsv2], continuing command...
There are pending nodes to be drained:
 worker-1
...
refusalwhyflag
DaemonSet podsthe DaemonSet would recreate them on the same node instantly--ignore-daemonsets (leaves them running)
pods with emptyDirits data is deleted with the pod--delete-emptydir-data
pods with no controllernothing will recreate them - evicting = deleting forever--force
a PDB would be violatedbudget says not nownone: drain waits and retries (fix the budget or add replicas)

The standard maintenance command is therefore kubectl drain worker-1 --ignore-daemonsets --delete-emptydir-data, and --force is a decision, not a habit: find out what that bare pod is first. Mirror pods (static pods) are skipped - the apiserver cannot stop them anyway.

uncordon undoes the cordon. It does not move anything back: the pods that left stay where they went until something reschedules them (a rollout, a scale, another drain). A freshly uncordoned node is empty - expected. Give worker-1 back now (the drain above stopped before evicting anything):

$ kubectl uncordon worker-1
node/worker-1 uncordoned

Unplanned: the node stops answering

The kubelet posts status every 10s and renews its Lease (a tiny heartbeat object, one per node) in the kube-node-lease namespace. The node controller (in kube-controller-manager) watches those:

time since last heartbeatwhat you see
< 50sReady (the default --node-monitor-grace-period - how long the node controller waits for a heartbeat - is 50s since 1.32; it was 40s)
> 50sReady Unknown: Kubelet stopped posting node status. - STATUS NotReady
immediately thentaints node.kubernetes.io/unreachable:NoSchedule and :NoExecute
+300spods without a longer toleration are evicted (marked for deletion)

Two flavours of NotReady, and they point to different places:

# an illustration: worker-2 stopped answering a minute ago (the node incidents)
$ kubectl describe node worker-2 | grep -A6 Conditions
Conditions:
  Type             Status    LastHeartbeatTime                 LastTransitionTime                Reason              Message
  ----             ------    -----------------                 ------------------                ------              -------
  MemoryPressure   Unknown   Tue, 22 Sep 2026 20:00:00 +0000   Tue, 22 Sep 2026 20:00:51 +0000   NodeStatusUnknown   Kubelet stopped posting node status.
  DiskPressure     Unknown   ...
  Ready            Unknown   ...

Taint-based eviction: the 300 seconds

Every pod gets two tolerations added by the DefaultTolerationSeconds admission plugin:

# an illustration
kubectl get pod web-7c9 -o jsonpath='{.spec.tolerations}' | jq .
[
  { "effect": "NoExecute", "key": "node.kubernetes.io/not-ready", "operator": "Exists", "tolerationSeconds": 300 },
  { "effect": "NoExecute", "key": "node.kubernetes.io/unreachable", "operator": "Exists", "tolerationSeconds": 300 }
]

So pods on a dead node are left alone for 5 minutes (in case the node comes back), then evicted and recreated elsewhere by their controllers. Consequences:

# an illustration: six minutes after worker-2 died
$ kubectl get pods -o wide
NAME                   READY   STATUS        RESTARTS   AGE     NODE
web-w9m2pkd2r6-lng4v   1/1     Running       0          12s     worker-1
web-w9m2pkd2r6-rlfg7   1/1     Terminating   0          6m19s   worker-2

The Terminating one is the ghost on the dead node; the Running one is its replacement. When the node returns, its kubelet sees the deletion, kills the container and the ghost disappears.

Pods show Running on a NotReady node

Until eviction, pods on an unreachable node still show 1/1 Running: that is the last status the kubelet reported, not current truth. Nobody can know whether they run. The Pod's Ready condition goes False, so Services stop sending traffic to them - which is why a NotReady node does not usually mean errors for users, only less capacity.

Planned reboot: drain first

sudo reboot on a node without draining is an unplanned outage in slow motion: everything on it dies at once, and nothing is recreated elsewhere for 5+ minutes (it may come back first). Drain, reboot, wait for Ready, uncordon.

What you can now do

Why it helps

Node maintenance is weekly platform work: OS patching, kernel updates, upgrades, hardware moves. Doing it without user impact means drain with the right flags, respecting PDBs, and remembering to uncordon. Knowing the three refusals (DaemonSets, emptyDir, bare pods) and what --force really does keeps you from deleting someone's debug pod forever, or someone's data.

The unplanned side explains incident timelines: why a dead node costs up to about 50 plus 300 seconds of reduced capacity, why pods show Running on a NotReady node (last reported status), and why a StatefulSet pod stays Terminating forever and needs a human decision. That last one is a classic senior interview question and a real 3am call.

Commands in this lesson

kubectl

FAQ

What's the difference between cordon and drain?

Cordon only marks the node unschedulable (spec.unschedulable: true plus the unschedulable taint): nothing new lands there, nothing running moves. Drain is cordon plus evicting every evictable pod through the Eviction API, respecting PodDisruptionBudgets and waiting for each pod to terminate gracefully. Uncordon removes the mark, but doesn't bring pods back.

Why does drain refuse some pods?

Three kinds, each for a reason. DaemonSet pods would be recreated on the same node immediately (--ignore-daemonsets leaves them). Pods with emptyDir lose that data (--delete-emptydir-data). Pods without a controller would be gone forever (--force). A PDB that would be violated makes drain wait and retry; no flag fixes that except --disable-eviction, as a last resort.

Why do pods show Running on a node that's NotReady?

That's the last status the kubelet reported before it stopped reporting; nobody knows the current truth. The pods' Ready condition goes False, so Services stop sending them traffic. That's why a NotReady node usually means less capacity rather than user errors, until the pods are evicted and recreated elsewhere.

Why is my StatefulSet pod stuck Terminating after its node died?

The eviction marked it for deletion, but only the node's kubelet can confirm its containers stopped, and that kubelet is gone. The StatefulSet won't create db-0 elsewhere while db-0 still exists, to guarantee at most one. Bring the node back, or if it's truly dead, delete the Node object or force-delete the pod after confirming the machine is off.

Can I just reboot a node without draining?

You can, but it's an unplanned outage in slow motion: everything on the node dies at once, and nothing is recreated elsewhere for several minutes (the node may even come back first). For a planned reboot: drain, reboot, wait for Ready, uncordon. That way replacements are running elsewhere before anything stops.

In an interview Junior

How do you safely take a node out for maintenance, and what happens when a node dies unexpectedly?

Planned: kubectl drain worker-1 --ignore-daemonsets --delete-emptydir-data - it cordons the node (no new pods) and evicts every pod through the Eviction API, respecting PDBs. --force (bare pods, deleted for good) is a decision, not a habit. Do the work, wait for Ready, kubectl uncordon - nothing moves back by itself.

Unplanned, minute by minute:

Also asked: What is the difference between Ready False and Ready Unknown on a node? · Why does a StatefulSet pod not move off a dead node? · What does kubectl uncordon do, and what does it not do?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.