Rules about other pods
The problem. You run 3 replicas "for high availability", a node dies, and all 3 die with it - the scheduler had put them on the same node. Or the opposite: a cache is useless unless it sits on the same node as the API it serves. Both are rules about other pods, not about nodes.
What you need to know already: node labels and node affinity (17.11), taints (17.13), ReplicaSets and rolling updates (15.16).
Node affinity looks at node labels. Inter-pod rules look at which pods already run where, grouped by a node label called the topologyKey:
affinity:
podAntiAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
- labelSelector:
matchLabels: {app: web}
topologyKey: kubernetes.io/hostname
labelSelector picks which pods count (here: pods labelled app=web). Read it as: "do not put me in a topology domain (a group of nodes that share the same value of the topologyKey label - here: a single node, because every node has a unique kubernetes.io/hostname) that already runs a pod matching app=web". With topologyKey: topology.kubernetes.io/zone the domain is a zone.
- podAffinity = "put me in a domain that has a matching pod" (co-locate a cache with its API, a sidecar-like helper with its main app).
- podAntiAffinity = "put me in a domain that has no matching pod" (spread replicas so one node or zone failure cannot take them all).
required...= hard filter,preferred...= weighted score, as with node affinity.- The selector matches pods in the pod's own namespace unless
namespacesornamespaceSelectorsays otherwise.
The classic: all three replicas landed on one node
The default scheduler already tries to spread a ReplicaSet's pods (a built-in soft spread score), but it is only a preference: resources, image locality and other scores can outweigh it, and you end up with three replicas on one node that then dies. Two fixes.
Fix 1 - required anti-affinity per node. Hard: never two on one node.
# an illustration: web with the spread constraint above
kubectl get pods -o wide -l app=web
NAME READY STATUS NODE
web-7c9f4d8b6c-4kq2x 1/1 Running worker-1
web-7c9f4d8b6c-9zt7m 1/1 Running worker-2
web-7c9f4d8b6c-mx2qp 0/1 Pending <none>
Warning FailedScheduling default-scheduler 0/3 nodes are available: 1 node(s) had untolerated taint {node-role.kubernetes.io/control-plane: }, 2 node(s) didn't match pod anti-affinity rules. preemption: 0/3 nodes are available: 1 Preemption is not helpful for scheduling, 2 No preemption victims found for incoming pod.
Three replicas, two workers: the third can never run. That is the cost of "hard": replicas > domains = Pending. It also blocks rolling updates with maxSurge if there is no spare domain for the surge pod. The other wording you will see is "node(s) didn't satisfy existing pods anti-affinity rules" - that is the running pods' anti-affinity rejecting the newcomer (anti-affinity is enforced both ways).
Fix 2 - topology spread constraints (topologySpreadConstraints: rules that say "keep my pods evenly spread over these domains"). Built for exactly this:
spec:
topologySpreadConstraints:
- maxSkew: 1
topologyKey: kubernetes.io/hostname
whenUnsatisfiable: DoNotSchedule # or ScheduleAnyway
labelSelector:
matchLabels: {app: web}
- maxSkew: the maximum allowed difference (the skew) between the domain with the most matching pods and the one with the fewest.
1= as even as possible. - whenUnsatisfiable:
DoNotSchedule= hard (Pending, "node(s) didn't match pod topology spread constraints"),ScheduleAnyway= soft (prefer the emptier domain). - labelSelector: which pods are counted - normally your own app's pods.
- minDomains: with DoNotSchedule, "act as if the global minimum were 0 until there are at least N domains" - forces spreading even while zones are empty.
- nodeAffinityPolicy (default Honor) / nodeTaintsPolicy (default Ignore): which nodes count as domains at all.
With 3 replicas, 2 workers, maxSkew 1: 2 + 1 is allowed (skew 1), so all three run. With 5 replicas: 3 + 2.
One trap on kubeadm clusters: with the default nodeTaintsPolicy: Ignore, the tainted control-plane node still counts as a domain - it has a hostname label - with zero matching pods, forever. The global minimum is then 0, and maxSkew 1 allows only one pod per worker: replica 3 is Pending with "node(s) didn't match pod topology spread constraints". nodeTaintsPolicy: Honor removes nodes the pod does not tolerate from the calculation (GA since 1.33). On a cloud cluster the same thing happens with a tainted system node pool.
That is the answer to "the difference between them":
Anti-affinity is binary - "at most one per domain" - so replicas beyond the number of domains cannot run. Topology spread is arithmetic - "never more than maxSkew apart" - so it spreads evenly and still runs any number of replicas.
Nodes that lack the topologyKey label are not a domain; with DoNotSchedule such a node is filtered: "node(s) didn't match pod topology spread constraints (missing required label)".
Zones, and combining keys
topologySpreadConstraints:
- maxSkew: 1
topologyKey: topology.kubernetes.io/zone # spread across zones first
whenUnsatisfiable: DoNotSchedule
labelSelector: {matchLabels: {app: web}}
- maxSkew: 1
topologyKey: kubernetes.io/hostname # and across nodes within a zone
whenUnsatisfiable: ScheduleAnyway
labelSelector: {matchLabels: {app: web}}
Every constraint must hold (they are ANDed). A common production default: hard across zones, soft across nodes.
Pod affinity: co-location
# the cache pods should run on the same node as an api pod
affinity:
podAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
- labelSelector: {matchLabels: {app: api}}
topologyKey: kubernetes.io/hostname
If no api pod runs anywhere yet, the cache pod is Pending with "node(s) didn't match pod affinity rules" - unless it matches its own selector (the first pod of a self-affine group may go anywhere). Use preferred affinity when you can: hard co-location couples the two workloads' scheduling fates.
Costs to know
Inter-pod affinity is the most expensive filter in the scheduler (it considers every matching pod on every node); upstream warns against it in clusters of several hundred nodes. Topology spread is cheaper and is the recommended tool for spreading. Both are IgnoredDuringExecution: nothing rebalances pods that are already running when a new node appears - that needs a restart or the descheduler (an optional add-on that evicts pods so they can be placed better).
What you can now do
- Spread replicas with anti-affinity (at most one per domain) or topology spread (even, any count).
- Co-locate pods with pod affinity, and predict the Pending message when a rule cannot be met.
- Explain why a tainted control-plane node can break a hostname spread (
nodeTaintsPolicy).