The opposite direction
The problem. You bought one big node for batch jobs. Node affinity can send the batch pods there - but nothing stops every other pod from landing on it too and eating its capacity. You need a way for the node to say "keep out".
What you need to know already: node labels and affinity (17.11), eviction (17.3), DaemonSets (15.22).
Affinity is a property of the pod that pulls it towards nodes. A taint is a property of the node that pushes pods away - every pod, unless the pod carries a matching toleration.
$ kubectl describe node cp-1 | grep Taints
Taints: node-role.kubernetes.io/control-plane:NoSchedule
That one taint is why none of your pods ever ran on cp-1: kubeadm taints the control plane, and ordinary pods do not tolerate it. The control plane's own pods (apiserver, scheduler...) and the DaemonSets (calico-node - the network plugin, kube-proxy) do.
Read the taint as three parts: key node-role.kubernetes.io/control-plane, an empty value, and the effect NoSchedule (what happens to pods that do not tolerate it).
Adding and removing taints
$ kubectl taint node worker-2 dedicated=batch:NoSchedule
node/worker-2 tainted
$ kubectl taint node worker-2 dedicated=batch:NoSchedule
error: node worker-2 already has dedicated taint(s) with same effect(s) and --overwrite is false
$ kubectl taint node worker-2 dedicated:NoSchedule-
node/worker-2 untainted
kubectl taint node NAME key=value:Effect adds a taint. Format: key[=value]:Effect, and a trailing - removes. A node may carry several taints; a pod must tolerate all of them (all the NoSchedule/NoExecute ones) to land there.
The three effects
| effect | new pods | pods already running |
|---|---|---|
NoSchedule | not scheduled unless tolerated | left alone |
PreferNoSchedule | the scheduler tries to avoid the node (a score) | left alone |
NoExecute | not scheduled unless tolerated | evicted unless tolerated (optionally after tolerationSeconds) |
NoExecute is the only placement mechanism that acts on running pods. Below, kubectl get events lists recent events; --field-selector involvedObject.kind=Pod keeps only events about pods:
# an illustration from the taints mission
$ kubectl taint node worker-1 maintenance=true:NoExecute
node/worker-1 tainted
$ kubectl get events -n shop --field-selector involvedObject.kind=Pod | grep Taint
2s Normal TaintManagerEviction pod/web-5d8f-2kq9x Marking for deletion Pod shop/web-5d8f-2kq9x
Tolerations
spec:
tolerations:
- key: dedicated # Equal (the default): key, value and effect must match
operator: Equal
value: batch
effect: NoSchedule
- key: gpu # Exists: any value
operator: Exists
effect: NoSchedule
- operator: Exists # no key + Exists = tolerate EVERYTHING (DaemonSets for system agents)
Rules: an empty effect matches all effects; operator: Exists must not have a value; a toleration without key must use Exists.
tolerationSeconds (NoExecute only) = "I tolerate this for N seconds, then evict me". Every pod gets two of these for free from admission:
Tolerations: node.kubernetes.io/not-ready:NoExecute op=Exists for 300s
node.kubernetes.io/unreachable:NoExecute op=Exists for 300s
That is the famous five minutes: when a node goes NotReady/unreachable (its kubelet stops reporting in), the node controller (a loop in the control plane that watches node health) adds those NoExecute taints and your pods are evicted 300s later - not before. Latency-critical apps shorten it (tolerationSeconds: 30); stateful ones sometimes lengthen it.
Built-in taints
The node lifecycle controller and the kubelet mirror node conditions as taints (all under node.kubernetes.io/):
not-ready:NoExecute unreachable:NoExecute unschedulable:NoSchedule (cordon)
memory-pressure:NoSchedule disk-pressure:NoSchedule pid-pressure:NoSchedule
network-unavailable:NoSchedule
kubectl cordon NODE marks a node "no new pods" before maintenance (lesson 17.26 uses it). It is literally spec.unschedulable: true + the unschedulable:NoSchedule taint. DaemonSet pods automatically tolerate the pressure and unschedulable taints, which is why they still start on a cordoned node.
Tolerating is permission, not attraction
The most common taint mistake: a toleration allows a pod on a tainted node; it does not send it there. A batch pod that tolerates dedicated=batch is just as happy on worker-1.
Dedicated nodes therefore need both halves:
# on the node: kubectl taint node worker-2 dedicated=batch:NoSchedule
# kubectl label node worker-2 dedicated=batch
spec:
tolerations: # may go to the batch node...
- {key: dedicated, operator: Equal, value: batch, effect: NoSchedule}
nodeSelector: # ...and must
dedicated: batch
- the taint keeps everybody else off the node,
- the toleration lets the batch pods on,
- the selector/affinity keeps the batch pods on it.
GPU nodes, licensed-software nodes, nodes in a locked-down network segment for card payments (PCI-DSS, the card industry's security standard): same pattern. Managed clusters use it to keep system add-ons on their own nodes (CriticalAddonsOnly=true:NoSchedule).
In FailedScheduling
0/3 nodes are available: 1 node(s) had untolerated taint {dedicated: batch}, 1 node(s) had untolerated taint {node-role.kubernetes.io/control-plane: }, 1 Insufficient cpu.
{key: value} - the value is empty for the control-plane taint. Taints are checked before node affinity and resources, so a tainted node is reported as "untolerated taint" even if it also fails other rules.
What you can now do
- Taint a node, tolerate the taint, and build a dedicated node (taint + label + selector).
- Predict what each effect does to new and to running pods.
- Explain the default 300s
not-ready/unreachabletolerations.