OnCallReady

Lesson 17.11 · Kubernetes: Scheduling, Health & Security · 14 min read

nodeSelector and node affinity

In plain words

Imagine a sports day where each child carries a card with their preferences. Some cards say "I must be in a team with a red shirt" (a hard rule: no red team, no play). Some say "I'd like to be near the shade, if possible" (a wish that counts for something but can be ignored). The teams' shirts are the labels.

Kubernetes places pods by labels on nodes. nodeSelector is the simplest hard rule: every key must match exactly. Node affinity is the expressive version: required... rules are filters (with In, NotIn, Exists, OR between terms), preferred... rules add a weighted score. The IgnoredDuringExecution part means the card is only checked when the pod is placed; change the node's labels later and nothing moves.

Placement starts with node labels

The problem. Your database must run on the node with the fast SSD disk; the batch reports must stay off it. By default the scheduler puts a pod on any node with room. Placement rules let you say where a pod may - or must - go.

What you need to know already: labels and selectors (15.26), the scheduler (15.5), reading FailedScheduling events (17.1).

Every placement rule in this part matches labels on nodes (key=value tags, exactly like pod labels). A kubeadm node (kubeadm = the standard tool that builds a cluster; this lab was built with it) comes with a handful. --show-labels adds a LABELS column:

$ kubectl get nodes --show-labels
NAME       STATUS   ROLES           AGE   VERSION   LABELS
cp-1       Ready    control-plane   12d   v1.34.1   beta.kubernetes.io/arch=arm64,beta.kubernetes.io/os=linux,kubernetes.io/arch=arm64,kubernetes.io/hostname=cp-1,kubernetes.io/os=linux,node-role.kubernetes.io/control-plane=,node.kubernetes.io/exclude-from-external-load-balancers=
worker-1   Ready    <none>          12d   v1.34.1   beta.kubernetes.io/arch=arm64,beta.kubernetes.io/os=linux,kubernetes.io/arch=arm64,kubernetes.io/hostname=worker-1,kubernetes.io/os=linux
worker-2   Ready    <none>          12d   v1.34.1   beta.kubernetes.io/arch=arm64,beta.kubernetes.io/os=linux,kubernetes.io/arch=arm64,kubernetes.io/hostname=worker-2,kubernetes.io/os=linux

The built-in ones: kubernetes.io/arch (CPU type), kubernetes.io/os, kubernetes.io/hostname (the node's name), node-role.kubernetes.io/control-plane (marks control-plane nodes). The beta. ones are old duplicates.

-L key1,key2 shows chosen labels as columns, which is much easier to read:

$ kubectl get nodes -L kubernetes.io/arch,topology.kubernetes.io/zone,disktype
NAME       STATUS   ROLES           AGE   VERSION   ARCH    ZONE   DISKTYPE
cp-1       Ready    control-plane   12d   v1.34.1   arm64
worker-1   Ready    <none>          12d   v1.34.1   arm64
worker-2   Ready    <none>          12d   v1.34.1   arm64

Cloud clusters add topology.kubernetes.io/zone (the zone: one data centre building of a cloud region), topology.kubernetes.io/region and node.kubernetes.io/instance-type (the machine size) automatically. On kubeadm you add your own with kubectl label node NAME key=value (--overwrite to change an existing value, key- with a trailing dash to remove it):

$ kubectl label node worker-1 disktype=ssd
node/worker-1 labeled
$ kubectl label node worker-1 disktype=nvme
error: 'disktype' already has a value (ssd), and --overwrite is false
$ kubectl label node worker-1 disktype=nvme --overwrite
node/worker-1 labeled
$ kubectl label node worker-1 disktype-
node/worker-1 unlabeled

nodeSelector: the simple, hard rule

nodeSelector sits in the pod spec (in a Deployment: under spec.template.spec) and lists labels the node must have:

spec:
  nodeSelector:
    disktype: ssd

Every key must match exactly (AND). Nothing matching = the pod stays Pending:

Warning  FailedScheduling  default-scheduler  0/3 nodes are available: 1 node(s) had untolerated taint {node-role.kubernetes.io/control-plane: }, 2 node(s) didn't match Pod's node affinity/selector. preemption: 0/3 nodes are available: 3 Preemption is not helpful for scheduling.

Read the tally: cp-1 was rejected by its taint (a "keep out" mark on the node, next lesson - checked first), the two workers by the selector. "Preemption is not helpful" - evicting pods cannot change a node's labels, so the scheduler does not even try.

Node affinity: the expressive version

Affinity = "attraction": a rule that pulls a pod towards certain nodes. Node affinity does what nodeSelector does, with more options (OR, NOT, "prefer"). The long field names read as sentences: required during scheduling, ignored during execution = a hard rule, checked only when the pod is placed.

spec:
  affinity:
    nodeAffinity:
      requiredDuringSchedulingIgnoredDuringExecution:     # HARD
        nodeSelectorTerms:
        - matchExpressions:
          - key: disktype
            operator: In
            values: [ssd, nvme]
          - key: node.kubernetes.io/instance-type
            operator: NotIn
            values: [Standard_B2s]          # a small machine size
      preferredDuringSchedulingIgnoredDuringExecution:    # SOFT
      - weight: 80
        preference:
          matchExpressions:
          - key: topology.kubernetes.io/zone
            operator: In
            values: [zone-a]

The logic, precisely:

IgnoredDuringExecution: the half of the name people skip

requiredDuringScheduling**IgnoredDuringExecution** means the rule is checked once, at scheduling time. Change the node's labels afterwards and nothing happens to pods already running there:

# an illustration: a Deployment with the affinity above (the affinity mission)
$ kubectl get pods -o wide
NAME                  READY   STATUS    NODE
db-6f9c7b8d4d-2xk9p   1/1     Running   worker-1
$ kubectl label node worker-1 disktype-
node/worker-1 unlabeled
$ kubectl get pods -o wide
NAME                  READY   STATUS    NODE
db-6f9c7b8d4d-2xk9p   1/1     Running   worker-1        <- still there
$ kubectl delete pod db-6f9c7b8d4d-2xk9p
$ kubectl get pods
NAME                  READY   STATUS    NODE
db-6f9c7b8d4d-q8m2z   0/1     Pending   <none>          <- the replacement cannot be placed

The danger is latent: your app looks fine until the next restart, node drain or rollout, and then nothing can be placed. (A ...RequiredDuringExecution variant that would evict has been proposed for years; it does not exist. Taints with NoExecute - next lesson - are the way to push running pods off a node.)

nodeName: bypassing the scheduler

spec.nodeName: worker-2 skips the scheduler entirely - the kubelet on worker-2 just runs the pod. No taint checks, no resource checks (the kubelet will reject it with OutOfcpu if it really does not fit). It is what the scheduler itself writes; as a human, use it only for debugging a specific node.

Where to use what

What you can now do

Why it helps

You'll use this for real placement needs: arm64 versus amd64 images, GPU pools, nodes in a PCI-DSS network segment, spot versus on-demand pools, "prefer the local zone". Every one of them is a node label plus a selector or affinity rule.

The latent-danger part is what bites in production: someone removes a label from a node, everything keeps running, and the next drain or rollout leaves a critical pod Pending with "didn't match Pod's node affinity/selector". Knowing IgnoredDuringExecution lets you predict that during a change review instead of discovering it at 3am. On the CKA, labelling a node and making a pod land there is a near-guaranteed task.

Commands in this lesson

kubectl

FAQ

What's the difference between nodeSelector and node affinity?

nodeSelector is a simple map: every key must equal the given value (AND). Required node affinity does the same job with more expressive operators (In, NotIn, Exists, DoesNotExist, Gt, Lt) and OR between terms. Preferred affinity adds soft preferences with weights. If you use both nodeSelector and required affinity, both must hold.

If I remove a label from a node, do pods using it move?

No. requiredDuringSchedulingIgnoredDuringExecution is checked only at scheduling. Running pods stay; the problem appears when they're next recreated (a rollout, a drain, a crash) and the replacement can't be placed. To push running pods off a node, use a NoExecute taint or drain it.

How do terms and expressions combine in node affinity?

Multiple nodeSelectorTerms are ORed: the node must match any one term. The matchExpressions inside one term are ANDed: all must match. So two terms mean "either of these", and one term with two expressions means "both of these". Getting that wrong is a common reason a pod lands somewhere unexpected.

Does preferred affinity guarantee the pod lands on a matching node?

No. Preferred rules add their weight (1-100) to the score of each matching node, and the scheduler usually picks the highest total, but resources, spreading and other scores also count. If no node matches, the pod is scheduled anyway, somewhere. Use it for wishes, not for requirements.

Does node affinity keep other pods off my special nodes?

No. Affinity attracts the pods that declare it; it doesn't repel anyone else. Any pod without a conflicting rule can still land on your GPU node. To reserve nodes for a workload, taint the nodes so other pods are repelled, and give your workload a toleration plus the affinity. That's the next lesson.

In an interview Junior

How do you make a pod run only on specific nodes?

Label the nodes, then select them:

Nothing matching leaves the pod Pending: "didn't match Pod's node affinity/selector".

Two things to say: IgnoredDuringExecution means the rule is checked only at scheduling - remove a label later and running pods stay, until the next restart cannot be placed. And affinity only attracts; it does not keep other pods off those nodes - that needs a taint.

Also asked: What does IgnoredDuringExecution mean, and what risk does it create? · How would you run a workload only on arm64 nodes in a mixed cluster? · What does setting spec.nodeName do?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.