OnCallReady

Lesson 17.13 · Kubernetes: Scheduling, Health & Security · 14 min read

Taints and tolerations: the node repels, the pod tolerates

In plain words

Imagine some rooms in a school have a sign on the door: "Staff only". Everyone stays out, unless they have a staff badge. The badge doesn't make you go into that room; it only means you're allowed in if you happen to be sent there. Some signs are stricter: "Closed for cleaning, everyone out now", and even people already inside must leave (unless their badge says they can stay a few minutes).

A taint is the sign on the node; a toleration is the badge on the pod. NoSchedule keeps new pods out, PreferNoSchedule is a polite "please don't", NoExecute also evicts pods already running. Because a toleration only permits, dedicated nodes need a taint (keep others off) plus a nodeSelector or affinity (keep your pods on).

The opposite direction

The problem. You bought one big node for batch jobs. Node affinity can send the batch pods there - but nothing stops every other pod from landing on it too and eating its capacity. You need a way for the node to say "keep out".

What you need to know already: node labels and affinity (17.11), eviction (17.3), DaemonSets (15.22).

Affinity is a property of the pod that pulls it towards nodes. A taint is a property of the node that pushes pods away - every pod, unless the pod carries a matching toleration.

$ kubectl describe node cp-1 | grep Taints
Taints:             node-role.kubernetes.io/control-plane:NoSchedule

That one taint is why none of your pods ever ran on cp-1: kubeadm taints the control plane, and ordinary pods do not tolerate it. The control plane's own pods (apiserver, scheduler...) and the DaemonSets (calico-node - the network plugin, kube-proxy) do.

Read the taint as three parts: key node-role.kubernetes.io/control-plane, an empty value, and the effect NoSchedule (what happens to pods that do not tolerate it).

Adding and removing taints

$ kubectl taint node worker-2 dedicated=batch:NoSchedule
node/worker-2 tainted
$ kubectl taint node worker-2 dedicated=batch:NoSchedule
error: node worker-2 already has dedicated taint(s) with same effect(s) and --overwrite is false
$ kubectl taint node worker-2 dedicated:NoSchedule-
node/worker-2 untainted

kubectl taint node NAME key=value:Effect adds a taint. Format: key[=value]:Effect, and a trailing - removes. A node may carry several taints; a pod must tolerate all of them (all the NoSchedule/NoExecute ones) to land there.

The three effects

effectnew podspods already running
NoSchedulenot scheduled unless toleratedleft alone
PreferNoSchedulethe scheduler tries to avoid the node (a score)left alone
NoExecutenot scheduled unless toleratedevicted unless tolerated (optionally after tolerationSeconds)

NoExecute is the only placement mechanism that acts on running pods. Below, kubectl get events lists recent events; --field-selector involvedObject.kind=Pod keeps only events about pods:

# an illustration from the taints mission
$ kubectl taint node worker-1 maintenance=true:NoExecute
node/worker-1 tainted
$ kubectl get events -n shop --field-selector involvedObject.kind=Pod | grep Taint
2s   Normal   TaintManagerEviction   pod/web-5d8f-2kq9x   Marking for deletion Pod shop/web-5d8f-2kq9x

Tolerations

spec:
  tolerations:
  - key: dedicated          # Equal (the default): key, value and effect must match
    operator: Equal
    value: batch
    effect: NoSchedule
  - key: gpu                # Exists: any value
    operator: Exists
    effect: NoSchedule
  - operator: Exists        # no key + Exists = tolerate EVERYTHING (DaemonSets for system agents)

Rules: an empty effect matches all effects; operator: Exists must not have a value; a toleration without key must use Exists.

tolerationSeconds (NoExecute only) = "I tolerate this for N seconds, then evict me". Every pod gets two of these for free from admission:

Tolerations:                 node.kubernetes.io/not-ready:NoExecute op=Exists for 300s
                             node.kubernetes.io/unreachable:NoExecute op=Exists for 300s

That is the famous five minutes: when a node goes NotReady/unreachable (its kubelet stops reporting in), the node controller (a loop in the control plane that watches node health) adds those NoExecute taints and your pods are evicted 300s later - not before. Latency-critical apps shorten it (tolerationSeconds: 30); stateful ones sometimes lengthen it.

Built-in taints

The node lifecycle controller and the kubelet mirror node conditions as taints (all under node.kubernetes.io/):

not-ready:NoExecute        unreachable:NoExecute      unschedulable:NoSchedule  (cordon)
memory-pressure:NoSchedule disk-pressure:NoSchedule   pid-pressure:NoSchedule
network-unavailable:NoSchedule

kubectl cordon NODE marks a node "no new pods" before maintenance (lesson 17.26 uses it). It is literally spec.unschedulable: true + the unschedulable:NoSchedule taint. DaemonSet pods automatically tolerate the pressure and unschedulable taints, which is why they still start on a cordoned node.

Tolerating is permission, not attraction

The most common taint mistake: a toleration allows a pod on a tainted node; it does not send it there. A batch pod that tolerates dedicated=batch is just as happy on worker-1.

Dedicated nodes therefore need both halves:

# on the node:   kubectl taint node worker-2 dedicated=batch:NoSchedule
#                kubectl label node worker-2 dedicated=batch
spec:
  tolerations:                        # may go to the batch node...
  - {key: dedicated, operator: Equal, value: batch, effect: NoSchedule}
  nodeSelector:                       # ...and must
    dedicated: batch

GPU nodes, licensed-software nodes, nodes in a locked-down network segment for card payments (PCI-DSS, the card industry's security standard): same pattern. Managed clusters use it to keep system add-ons on their own nodes (CriticalAddonsOnly=true:NoSchedule).

In FailedScheduling

0/3 nodes are available: 1 node(s) had untolerated taint {dedicated: batch}, 1 node(s) had untolerated taint {node-role.kubernetes.io/control-plane: }, 1 Insufficient cpu.

{key: value} - the value is empty for the control-plane taint. Taints are checked before node affinity and resources, so a tainted node is reported as "untolerated taint" even if it also fails other rules.

What you can now do

Why it helps

Taints explain things you'll see from day one: why nothing runs on the control plane (node-role.kubernetes.io/control-plane:NoSchedule), why pods take exactly five minutes to leave a dead node (the default tolerationSeconds: 300 for not-ready and unreachable), and why DaemonSets still start on cordoned nodes.

On a platform team you'll build dedicated pools with them: GPU nodes, licensed-software nodes, a card-payment (PCI-DSS) segment, system versus user pools on managed clusters (CriticalAddonsOnly). The classic bug you'll catch in review is a toleration without affinity: the batch pod may run on the batch nodes but happily lands anywhere. And you'll tune the eviction delay for latency-critical apps. Adding and removing taints is standard CKA material.

Commands in this lesson

kubectl

FAQ

If my pod tolerates a taint, will it be scheduled on that node?

Not necessarily. A toleration is permission, not attraction: it removes the obstacle, but the scheduler may still choose any other node. To keep a workload on dedicated nodes, combine the toleration with a nodeSelector or node affinity on a label those nodes carry. The taint keeps everyone else off.

Why do pods wait five minutes before leaving a failed node?

Admission adds two tolerations to every pod: node.kubernetes.io/not-ready:NoExecute and node.kubernetes.io/unreachable:NoExecute, each for 300 seconds. When a node goes NotReady, the node controller adds those taints, and the pods are evicted only when the 300 seconds run out. Set a shorter tolerationSeconds for latency-critical apps.

What's the difference between cordon and a NoSchedule taint?

Cordon is one: kubectl cordon sets spec.unschedulable: true, and the node gets the node.kubernetes.io/unschedulable:NoSchedule taint. It stops new pods but leaves running ones. DaemonSet pods tolerate it, which is why they still start there. kubectl drain is cordon plus evicting the pods.

How do I remove a taint?

Repeat it with a trailing dash: kubectl taint node worker-2 dedicated:NoSchedule- removes the dedicated taint with that effect, and kubectl taint node worker-2 dedicated- removes it for all effects. Adding the same key and effect again without --overwrite fails with "already has dedicated taint(s)".

What does a toleration with just operator: Exists do?

With no key and operator: Exists, it tolerates every taint. That's what system DaemonSets like CNI agents and log shippers use, because they must run on every node, including tainted, cordoned or pressured ones. For application pods it's almost always a mistake, since it lets them land on the control plane and any dedicated pool.

In an interview Junior

What are taints and tolerations?

A taint is on the node and repels pods: kubectl taint node worker-2 dedicated=batch:NoSchedule (key, value, effect; a trailing - removes it). A toleration is on the pod and permits it on a node with a matching taint.

The effects:

Examples you already have: the control plane's node-role.kubernetes.io/control-plane:NoSchedule, and the built-in not-ready/unreachable NoExecute taints that every pod tolerates for 300 s - the five minutes before pods leave a dead node.

The key point: a toleration is permission, not attraction. A dedicated node needs the taint (keeps others off), the toleration (lets yours on) and a nodeSelector or affinity (keeps yours there).

Also asked: How do you dedicate a set of nodes to one workload? · A node dies. What happens to its pods, and when? · What does kubectl cordon do to a node?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.