OnCallReady

Lesson 17.18 · Kubernetes: Scheduling, Health & Security · 15 min read

Reading FailedScheduling, and PriorityClass with preemption

In plain words

Imagine a teacher trying to find a seat for a new pupil in a school with three classrooms, and writing one note that explains every "no": "Room A is staff-only, Room B has no free chairs, Room C is for the choir." Then the teacher asks one more question: "could I move some less important pupils out to make room?" For the staff room or the choir room, moving pupils wouldn't help at all.

That's the FailedScheduling event: 0/3 nodes are available, then one reason per node (the first filter that rejected it), then a preemption: part. PriorityClass is the "importance" of pupils: higher priority pods are scheduled first, and a Pending high-priority pod may preempt (evict) lower-priority pods to make room, but only where that would actually help.

One sentence, every node

The problem. A pod is Pending and the event is one long sentence full of numbers. Read it right and it tells you, node by node, what to fix. And when the cluster is full, you need a way to say "this pod matters more than that one".

What you need to know already: requests (17.1), node affinity (17.11), taints (17.13), pod affinity and spread (17.15), eviction (17.3).

The scheduler works in two phases: filters (yes/no checks that throw nodes out) and scores (points for the nodes that are left; highest wins). It explains a failure in one event. Learn to read it as a table:

0/3 nodes are available: 1 node(s) had untolerated taint {node-role.kubernetes.io/control-plane: }, 1 Insufficient cpu, 1 node(s) didn't match Pod's node affinity/selector. preemption: 0/3 nodes are available: 1 No preemption victims found for incoming pod, 2 Preemption is not helpful for scheduling.

The filters run in a fixed order, and the first failure is what you see:

NodeUnschedulable      node(s) were unschedulable              (cordoned)
TaintToleration        node(s) had untolerated taint {k: v}
NodeAffinity           node(s) didn't match Pod's node affinity/selector
NodeResourcesFit       Insufficient cpu / Insufficient memory / Too many pods
VolumeBinding etc.     node(s) had volume node affinity conflict / unbound PVCs
PodTopologySpread      node(s) didn't match pod topology spread constraints
InterPodAffinity       node(s) didn't match pod affinity rules
                       node(s) didn't match pod anti-affinity rules
                       node(s) didn't satisfy existing pods anti-affinity rules

So a cordoned node that also lacks CPU shows up as "unschedulable" only. When you fix the first reason and the pod is still Pending, read the new message - the next filter may now be the one failing.

The message lives in two places. On the pod itself:

# an illustration: a Pending pod (the scheduling missions)
$ kubectl describe pod web-7c9f4-mx2qp | tail -3
$ kubectl get pod web-7c9f4-mx2qp -o jsonpath='{.status.conditions[?(@.type=="PodScheduled")].message}'

And as an event, which you can list for the whole cluster:

$ kubectl get events --field-selector reason=FailedScheduling -A

PriorityClass

A PriorityClass is a cluster-wide object that gives a name to a number; pods refer to it by name and get that number as their priority. Priority decides two things: the order of the scheduling queue (higher first), and preemption - a Pending pod may evict lower-priority pods to make room.

$ kubectl get priorityclass
NAME                      VALUE        GLOBAL-DEFAULT   AGE   PREEMPTIONPOLICY
system-cluster-critical   2000000000   false            12d   PreemptLowerPriority
system-node-critical      2000001000   false            12d   PreemptLowerPriority
$ kubectl create priorityclass payments-critical --value=100000 --description="the payment path"
priorityclass.scheduling.k8s.io/payments-critical created

--value = the priority number (bigger = more important), --description = a note for humans. There are also --global-default and --preemption-policy (below).

spec:
  priorityClassName: payments-critical     # admission resolves it into spec.priority

Preemption, step by step

  1. A high-priority pod is Unschedulable (Insufficient cpu).
  2. The scheduler finds a node where evicting some lower-priority pods would make it fit, choosing the victims that cost least (lowest priority, fewest).
  3. It evicts them gracefully (SIGTERM first, like a normal delete - 3.18) and records the plan on the preemptor (the pod that caused the preemption): status.nominatedNodeName: worker-1 (the NOMINATED NODE column of get pods -o wide).
  4. The victims get an event and a condition:
Normal  Preempted  default-scheduler  Preempted by pod 1f6ca233-ff7c-419a-a90e-88531839aa27 on node worker-1
  1. When they are gone, the preemptor is scheduled there; the room it made is reserved for it, so lower-priority pods cannot sneak in first.

Preemption ignores PodDisruptionBudgets (lesson 17.26) if it has no other choice (it tries to respect them), and it does not guarantee a spot - if the victims take long to terminate, something else may change. The victims' controllers recreate them, and the replacements are now the Pending ones.

The operational consequences

What you can now do

Why it helps

A Pending pod is the ticket you'll answer most often, and the scheduler has already written the diagnosis. Reading "1 node(s) had untolerated taint, 1 Insufficient cpu, 1 didn't match Pod's node affinity/selector" as a per-node table tells you exactly what to fix, and knowing the filter order explains why fixing one reason reveals the next.

PriorityClass has an operational edge: it's how coredns and kube-proxy win on a full cluster, and it's also how one team can accidentally evict another. When someone reports "my pods keep getting killed for no reason" and you find Preempted events, you'll know where to look. In platform design you'll pair priority classes with quotas so only the platform can use the high ones.

Commands in this lesson

kubectl

FAQ

Why does the message list only one reason per node?

The scheduler runs its filters in a fixed order and reports the first one that rejected each node. So a cordoned node that also lacks CPU shows up only as "unschedulable". The counts add up to the number of nodes. After you fix the first reason, read the new message, because the next filter may now be the one failing.

What does "Preemption is not helpful for scheduling" mean?

The scheduler asked whether evicting lower-priority pods from those nodes would let the pod fit, and the answer is no, because the rejection reason isn't about resources: taints, labels, affinity, unbound volumes. "No preemption victims found" means evicting could help in principle (resources), but no pod there has lower priority than this one.

What priority does a pod get if I don't set a class?

Zero, unless some PriorityClass has globalDefault: true, in which case pods without a class get that one's value. User-defined classes can go up to 1000000000; values above that are reserved for system-cluster-critical and system-node-critical, which coredns, kube-proxy and similar components use.

Does preemption respect PodDisruptionBudgets?

It tries to, preferring victims whose eviction doesn't violate a PDB, but if there's no other way it goes ahead anyway. PDBs are a best-effort consideration for preemption, not a guarantee. Victims are evicted gracefully and recreated by their controllers, and the replacements are then the Pending ones.

What is preemptionPolicy: Never for?

A class with preemptionPolicy: Never still puts its pods ahead in the scheduling queue, but they never evict anyone. It's for work that should go first when there's room, like important batch jobs, without hurting running workloads. The default is PreemptLowerPriority.

In an interview Junior

A pod is Pending. How do you find out why?

k describe pod POD and read the FailedScheduling event as a table:

0/3 nodes are available: 1 node(s) had untolerated taint {...control-plane: },
1 Insufficient cpu, 1 node(s) didn't match Pod's node affinity/selector.

Fix the reason for each node - and if the pod is still Pending, read the new message: the next filter in line may now be the one failing. No event at all? Check the pod has no nodeName and the scheduler is running. With a full cluster, a PriorityClass lets important pods be scheduled first and preempt lower-priority ones.

Also asked: What is a PriorityClass, and what is preemption? · You fixed the reason in FailedScheduling but the pod is still Pending. Why? · What does "Insufficient cpu" mean on a cluster with idle CPUs?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.