Placement starts with node labels
The problem. Your database must run on the node with the fast SSD disk; the batch reports must stay off it. By default the scheduler puts a pod on any node with room. Placement rules let you say where a pod may - or must - go.
What you need to know already: labels and selectors (15.26), the scheduler (15.5), reading FailedScheduling events (17.1).
Every placement rule in this part matches labels on nodes (key=value tags, exactly like pod labels). A kubeadm node (kubeadm = the standard tool that builds a cluster; this lab was built with it) comes with a handful. --show-labels adds a LABELS column:
$ kubectl get nodes --show-labels
NAME STATUS ROLES AGE VERSION LABELS
cp-1 Ready control-plane 12d v1.34.1 beta.kubernetes.io/arch=arm64,beta.kubernetes.io/os=linux,kubernetes.io/arch=arm64,kubernetes.io/hostname=cp-1,kubernetes.io/os=linux,node-role.kubernetes.io/control-plane=,node.kubernetes.io/exclude-from-external-load-balancers=
worker-1 Ready <none> 12d v1.34.1 beta.kubernetes.io/arch=arm64,beta.kubernetes.io/os=linux,kubernetes.io/arch=arm64,kubernetes.io/hostname=worker-1,kubernetes.io/os=linux
worker-2 Ready <none> 12d v1.34.1 beta.kubernetes.io/arch=arm64,beta.kubernetes.io/os=linux,kubernetes.io/arch=arm64,kubernetes.io/hostname=worker-2,kubernetes.io/os=linux
The built-in ones: kubernetes.io/arch (CPU type), kubernetes.io/os, kubernetes.io/hostname (the node's name), node-role.kubernetes.io/control-plane (marks control-plane nodes). The beta. ones are old duplicates.
-L key1,key2 shows chosen labels as columns, which is much easier to read:
$ kubectl get nodes -L kubernetes.io/arch,topology.kubernetes.io/zone,disktype
NAME STATUS ROLES AGE VERSION ARCH ZONE DISKTYPE
cp-1 Ready control-plane 12d v1.34.1 arm64
worker-1 Ready <none> 12d v1.34.1 arm64
worker-2 Ready <none> 12d v1.34.1 arm64
Cloud clusters add topology.kubernetes.io/zone (the zone: one data centre building of a cloud region), topology.kubernetes.io/region and node.kubernetes.io/instance-type (the machine size) automatically. On kubeadm you add your own with kubectl label node NAME key=value (--overwrite to change an existing value, key- with a trailing dash to remove it):
$ kubectl label node worker-1 disktype=ssd
node/worker-1 labeled
$ kubectl label node worker-1 disktype=nvme
error: 'disktype' already has a value (ssd), and --overwrite is false
$ kubectl label node worker-1 disktype=nvme --overwrite
node/worker-1 labeled
$ kubectl label node worker-1 disktype-
node/worker-1 unlabeled
nodeSelector: the simple, hard rule
nodeSelector sits in the pod spec (in a Deployment: under spec.template.spec) and lists labels the node must have:
spec:
nodeSelector:
disktype: ssd
Every key must match exactly (AND). Nothing matching = the pod stays Pending:
Warning FailedScheduling default-scheduler 0/3 nodes are available: 1 node(s) had untolerated taint {node-role.kubernetes.io/control-plane: }, 2 node(s) didn't match Pod's node affinity/selector. preemption: 0/3 nodes are available: 3 Preemption is not helpful for scheduling.
Read the tally: cp-1 was rejected by its taint (a "keep out" mark on the node, next lesson - checked first), the two workers by the selector. "Preemption is not helpful" - evicting pods cannot change a node's labels, so the scheduler does not even try.
Node affinity: the expressive version
Affinity = "attraction": a rule that pulls a pod towards certain nodes. Node affinity does what nodeSelector does, with more options (OR, NOT, "prefer"). The long field names read as sentences: required during scheduling, ignored during execution = a hard rule, checked only when the pod is placed.
spec:
affinity:
nodeAffinity:
requiredDuringSchedulingIgnoredDuringExecution: # HARD
nodeSelectorTerms:
- matchExpressions:
- key: disktype
operator: In
values: [ssd, nvme]
- key: node.kubernetes.io/instance-type
operator: NotIn
values: [Standard_B2s] # a small machine size
preferredDuringSchedulingIgnoredDuringExecution: # SOFT
- weight: 80
preference:
matchExpressions:
- key: topology.kubernetes.io/zone
operator: In
values: [zone-a]
The logic, precisely:
- Terms are ORed, the expressions inside one term are ANDed. Two terms = "match either". One term with two expressions = "match both".
- Operators:
In,NotIn,Exists,DoesNotExist,Gt,Lt(the last two compare integer label values:gpu-count Gt 1).NotInandDoesNotExistare how you express anti-affinity to nodes. - nodeSelector and required affinity together: both must hold.
- required is a filter. Not satisfied -> Pending, with the same "didn't match Pod's node affinity/selector" message.
- preferred is a score.
weight1-100 is added to every node that matches; the highest total usually wins, but resources, spreading and other scores still count. Nothing matches -> the pod is scheduled anyway, somewhere.
IgnoredDuringExecution: the half of the name people skip
requiredDuringScheduling**IgnoredDuringExecution** means the rule is checked once, at scheduling time. Change the node's labels afterwards and nothing happens to pods already running there:
# an illustration: a Deployment with the affinity above (the affinity mission)
$ kubectl get pods -o wide
NAME READY STATUS NODE
db-6f9c7b8d4d-2xk9p 1/1 Running worker-1
$ kubectl label node worker-1 disktype-
node/worker-1 unlabeled
$ kubectl get pods -o wide
NAME READY STATUS NODE
db-6f9c7b8d4d-2xk9p 1/1 Running worker-1 <- still there
$ kubectl delete pod db-6f9c7b8d4d-2xk9p
$ kubectl get pods
NAME READY STATUS NODE
db-6f9c7b8d4d-q8m2z 0/1 Pending <none> <- the replacement cannot be placed
The danger is latent: your app looks fine until the next restart, node drain or rollout, and then nothing can be placed. (A ...RequiredDuringExecution variant that would evict has been proposed for years; it does not exist. Taints with NoExecute - next lesson - are the way to push running pods off a node.)
nodeName: bypassing the scheduler
spec.nodeName: worker-2 skips the scheduler entirely - the kubelet on worker-2 just runs the pod. No taint checks, no resource checks (the kubelet will reject it with OutOfcpu if it really does not fit). It is what the scheduler itself writes; as a human, use it only for debugging a specific node.
Where to use what
nodeSelector- "must run on arm64", "must run on the GPU pool". Simple, readable.- required affinity - the same with OR, NotIn, Exists.
- preferred affinity - "prefer the local zone", "prefer spot nodes if available".
- Node affinity attracts pods to nodes; it does not keep other pods away. To reserve nodes for a workload you need taints as well - the next lesson.
What you can now do
- Label nodes and pin pods to them with
nodeSelectoror required node affinity. - Express "prefer" (weighted) and "never" (
NotIn) rules. - Explain why removing a label only bites at the next reschedule.