OnCallReady

Lesson 17.15 · Kubernetes: Scheduling, Health & Security · 13 min read

Pod affinity, anti-affinity and topology spread

In plain words

Imagine seating three twins' birthday guests at a party with three tables. One rule could be "never two guests from the same family at one table". Works with three tables; with four family members, one guest stands forever. A better rule is "keep the tables within one guest of each other": 2, 1, 1 is fine, so everyone sits and the family is still spread out. There's also the opposite wish: "seat the cake next to the kids".

Pod anti-affinity is the first rule: at most one matching pod per topology domain (a node, a zone), so extra replicas stay Pending. Topology spread constraints are the second: maxSkew limits the difference between domains, so any number of replicas runs and stays even. Pod affinity is the cake rule: co-locate with pods that match a selector.

Rules about other pods

The problem. You run 3 replicas "for high availability", a node dies, and all 3 die with it - the scheduler had put them on the same node. Or the opposite: a cache is useless unless it sits on the same node as the API it serves. Both are rules about other pods, not about nodes.

What you need to know already: node labels and node affinity (17.11), taints (17.13), ReplicaSets and rolling updates (15.16).

Node affinity looks at node labels. Inter-pod rules look at which pods already run where, grouped by a node label called the topologyKey:

affinity:
  podAntiAffinity:
    requiredDuringSchedulingIgnoredDuringExecution:
    - labelSelector:
        matchLabels: {app: web}
      topologyKey: kubernetes.io/hostname

labelSelector picks which pods count (here: pods labelled app=web). Read it as: "do not put me in a topology domain (a group of nodes that share the same value of the topologyKey label - here: a single node, because every node has a unique kubernetes.io/hostname) that already runs a pod matching app=web". With topologyKey: topology.kubernetes.io/zone the domain is a zone.

The classic: all three replicas landed on one node

The default scheduler already tries to spread a ReplicaSet's pods (a built-in soft spread score), but it is only a preference: resources, image locality and other scores can outweigh it, and you end up with three replicas on one node that then dies. Two fixes.

Fix 1 - required anti-affinity per node. Hard: never two on one node.

# an illustration: web with the spread constraint above
kubectl get pods -o wide -l app=web
NAME                   READY   STATUS    NODE
web-7c9f4d8b6c-4kq2x   1/1     Running   worker-1
web-7c9f4d8b6c-9zt7m   1/1     Running   worker-2
web-7c9f4d8b6c-mx2qp   0/1     Pending   <none>
Warning  FailedScheduling  default-scheduler  0/3 nodes are available: 1 node(s) had untolerated taint {node-role.kubernetes.io/control-plane: }, 2 node(s) didn't match pod anti-affinity rules. preemption: 0/3 nodes are available: 1 Preemption is not helpful for scheduling, 2 No preemption victims found for incoming pod.

Three replicas, two workers: the third can never run. That is the cost of "hard": replicas > domains = Pending. It also blocks rolling updates with maxSurge if there is no spare domain for the surge pod. The other wording you will see is "node(s) didn't satisfy existing pods anti-affinity rules" - that is the running pods' anti-affinity rejecting the newcomer (anti-affinity is enforced both ways).

Fix 2 - topology spread constraints (topologySpreadConstraints: rules that say "keep my pods evenly spread over these domains"). Built for exactly this:

spec:
  topologySpreadConstraints:
  - maxSkew: 1
    topologyKey: kubernetes.io/hostname
    whenUnsatisfiable: DoNotSchedule      # or ScheduleAnyway
    labelSelector:
      matchLabels: {app: web}

With 3 replicas, 2 workers, maxSkew 1: 2 + 1 is allowed (skew 1), so all three run. With 5 replicas: 3 + 2.

One trap on kubeadm clusters: with the default nodeTaintsPolicy: Ignore, the tainted control-plane node still counts as a domain - it has a hostname label - with zero matching pods, forever. The global minimum is then 0, and maxSkew 1 allows only one pod per worker: replica 3 is Pending with "node(s) didn't match pod topology spread constraints". nodeTaintsPolicy: Honor removes nodes the pod does not tolerate from the calculation (GA since 1.33). On a cloud cluster the same thing happens with a tainted system node pool.

That is the answer to "the difference between them":

Anti-affinity is binary - "at most one per domain" - so replicas beyond the number of domains cannot run. Topology spread is arithmetic - "never more than maxSkew apart" - so it spreads evenly and still runs any number of replicas.

Nodes that lack the topologyKey label are not a domain; with DoNotSchedule such a node is filtered: "node(s) didn't match pod topology spread constraints (missing required label)".

Zones, and combining keys

topologySpreadConstraints:
- maxSkew: 1
  topologyKey: topology.kubernetes.io/zone          # spread across zones first
  whenUnsatisfiable: DoNotSchedule
  labelSelector: {matchLabels: {app: web}}
- maxSkew: 1
  topologyKey: kubernetes.io/hostname              # and across nodes within a zone
  whenUnsatisfiable: ScheduleAnyway
  labelSelector: {matchLabels: {app: web}}

Every constraint must hold (they are ANDed). A common production default: hard across zones, soft across nodes.

Pod affinity: co-location

# the cache pods should run on the same node as an api pod
affinity:
  podAffinity:
    requiredDuringSchedulingIgnoredDuringExecution:
    - labelSelector: {matchLabels: {app: api}}
      topologyKey: kubernetes.io/hostname

If no api pod runs anywhere yet, the cache pod is Pending with "node(s) didn't match pod affinity rules" - unless it matches its own selector (the first pod of a self-affine group may go anywhere). Use preferred affinity when you can: hard co-location couples the two workloads' scheduling fates.

Costs to know

Inter-pod affinity is the most expensive filter in the scheduler (it considers every matching pod on every node); upstream warns against it in clusters of several hundred nodes. Topology spread is cheaper and is the recommended tool for spreading. Both are IgnoredDuringExecution: nothing rebalances pods that are already running when a new node appears - that needs a restart or the descheduler (an optional add-on that evicts pods so they can be placed better).

What you can now do

Why it helps

"All three replicas landed on the one node that died" is a real outage pattern, and the fix is one of these two tools. Knowing their difference lets you pick right: anti-affinity blocks rollouts and scale-ups once replicas outnumber domains, topology spread keeps spreading without blocking.

You'll also debug the subtle cases: a spread constraint that Pends the third replica on a kubeadm cluster because the tainted control-plane node counts as a domain with zero pods (nodeTaintsPolicy), or a required pod affinity that couples two teams' scheduling fates. Zone spreading is a standard requirement for anything highly available on a cloud cluster, and "anti-affinity versus topology spread" is a common interview question.

FAQ

Anti-affinity or topology spread for spreading replicas?

Usually topology spread. Anti-affinity is binary, "at most one per domain", so replicas beyond the number of domains are Pending and surge pods during rollouts can be blocked. Topology spread limits the difference between domains with maxSkew, so any number of replicas runs and stays balanced. It's also cheaper for the scheduler; upstream warns about inter-pod affinity in large clusters.

What is a topologyKey?

A node label that groups nodes into domains. kubernetes.io/hostname is unique per node, so each node is a domain; topology.kubernetes.io/zone groups nodes by zone. The rules are then evaluated per domain: "no matching pod in this zone", "no more than one pod difference between zones". Nodes without the label aren't part of any domain.

Why is my third replica Pending with maxSkew: 1 on two workers?

Probably the tainted control-plane node. With the default nodeTaintsPolicy: Ignore, it still counts as a hostname domain with zero matching pods forever, so the global minimum is 0 and maxSkew 1 allows only one pod per worker. Set nodeTaintsPolicy: Honor so nodes the pod doesn't tolerate are excluded. The same happens with tainted system pools in the cloud.

Do these rules rebalance pods when a new node joins?

No, they're all IgnoredDuringExecution. They only affect where new pods go. Running pods stay where they are, so after adding a node or recovering a zone, the spread stays uneven until pods are recreated (a rollout restart) or something like the descheduler evicts them.

Why is my pod with podAffinity Pending?

Required pod affinity needs a matching pod to already exist in the target domain. If no app: api pod runs anywhere yet, the cache pod can't be placed ("didn't match pod affinity rules"). The exception is a pod that matches its own selector, whose first replica may go anywhere. Prefer preferred affinity for co-location unless it's truly required.

In an interview Mid

How do you make sure replicas of a Deployment do not all end up on the same node or zone?

The default spreading is only a soft score, so state it:

A common production default: hard across zones, soft across nodes. Watch the kubeadm trap: with the default nodeTaintsPolicy: Ignore, the tainted control-plane node counts as an empty domain and blocks the spread - set nodeTaintsPolicy: Honor. And nothing rebalances running pods afterwards.

Also asked: What is the difference between pod affinity and pod anti-affinity? · What is maxSkew in a topology spread constraint? · How would you design pod placement for a service that must survive a zone outage?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.