One sentence, every node
The problem. A pod is Pending and the event is one long sentence full of numbers. Read it right and it tells you, node by node, what to fix. And when the cluster is full, you need a way to say "this pod matters more than that one".
What you need to know already: requests (17.1), node affinity (17.11), taints (17.13), pod affinity and spread (17.15), eviction (17.3).
The scheduler works in two phases: filters (yes/no checks that throw nodes out) and scores (points for the nodes that are left; highest wins). It explains a failure in one event. Learn to read it as a table:
0/3 nodes are available: 1 node(s) had untolerated taint {node-role.kubernetes.io/control-plane: }, 1 Insufficient cpu, 1 node(s) didn't match Pod's node affinity/selector. preemption: 0/3 nodes are available: 1 No preemption victims found for incoming pod, 2 Preemption is not helpful for scheduling.
0/3 nodes are available- three nodes were considered, none passed.- Then one reason per node (the first filter that rejected it), grouped and counted, sorted alphabetically: cp-1 -> taint, one worker -> not enough CPU (requests!), the other worker -> labels.
- The counts add up to the node total: each node is counted once, under the first filter that rejected it.
preemption:- the scheduler then asked "would evicting lower-priority pods help?". Not helpful for nodes rejected for reasons eviction cannot fix (taints, labels, affinity, unbound volumes). No preemption victims found for nodes where it could help in principle (resources) but no pod there has a lower priority than this one.
The filters run in a fixed order, and the first failure is what you see:
NodeUnschedulable node(s) were unschedulable (cordoned)
TaintToleration node(s) had untolerated taint {k: v}
NodeAffinity node(s) didn't match Pod's node affinity/selector
NodeResourcesFit Insufficient cpu / Insufficient memory / Too many pods
VolumeBinding etc. node(s) had volume node affinity conflict / unbound PVCs
PodTopologySpread node(s) didn't match pod topology spread constraints
InterPodAffinity node(s) didn't match pod affinity rules
node(s) didn't match pod anti-affinity rules
node(s) didn't satisfy existing pods anti-affinity rules
So a cordoned node that also lacks CPU shows up as "unschedulable" only. When you fix the first reason and the pod is still Pending, read the new message - the next filter may now be the one failing.
The message lives in two places. On the pod itself:
# an illustration: a Pending pod (the scheduling missions)
$ kubectl describe pod web-7c9f4-mx2qp | tail -3
$ kubectl get pod web-7c9f4-mx2qp -o jsonpath='{.status.conditions[?(@.type=="PodScheduled")].message}'
And as an event, which you can list for the whole cluster:
$ kubectl get events --field-selector reason=FailedScheduling -A
PriorityClass
A PriorityClass is a cluster-wide object that gives a name to a number; pods refer to it by name and get that number as their priority. Priority decides two things: the order of the scheduling queue (higher first), and preemption - a Pending pod may evict lower-priority pods to make room.
$ kubectl get priorityclass
NAME VALUE GLOBAL-DEFAULT AGE PREEMPTIONPOLICY
system-cluster-critical 2000000000 false 12d PreemptLowerPriority
system-node-critical 2000001000 false 12d PreemptLowerPriority
$ kubectl create priorityclass payments-critical --value=100000 --description="the payment path"
priorityclass.scheduling.k8s.io/payments-critical created
--value = the priority number (bigger = more important), --description = a note for humans. There are also --global-default and --preemption-policy (below).
spec:
priorityClassName: payments-critical # admission resolves it into spec.priority
- User classes may go up to 1000000000; above that is reserved for the two system classes (coredns and kube-proxy use them;
kube-systemis the namespace of the cluster's own components - checkkubectl get pods -n kube-system -o custom-columns=NAME:.metadata.name,PRIO:.spec.priority). - A pod without a class gets priority 0 - unless one class has
globalDefault: true. - A misspelled class is rejected at admission:
pods "x" is forbidden: no PriorityClass with name payment-critical was found(for a Deployment you find that on the ReplicaSet as FailedCreate). preemptionPolicy: Never= "jump the queue, but never evict anybody" - for batch work that should go first but not hurt others.
Preemption, step by step
- A high-priority pod is Unschedulable (
Insufficient cpu). - The scheduler finds a node where evicting some lower-priority pods would make it fit, choosing the victims that cost least (lowest priority, fewest).
- It evicts them gracefully (SIGTERM first, like a normal delete - 3.18) and records the plan on the preemptor (the pod that caused the preemption):
status.nominatedNodeName: worker-1(the NOMINATED NODE column ofget pods -o wide). - The victims get an event and a condition:
Normal Preempted default-scheduler Preempted by pod 1f6ca233-ff7c-419a-a90e-88531839aa27 on node worker-1
- When they are gone, the preemptor is scheduled there; the room it made is reserved for it, so lower-priority pods cannot sneak in first.
Preemption ignores PodDisruptionBudgets (lesson 17.26) if it has no other choice (it tries to respect them), and it does not guarantee a spot - if the victims take long to terminate, something else may change. The victims' controllers recreate them, and the replacements are now the Pending ones.
The operational consequences
- Priority without quotas is a way for one team to evict another. Pair classes with RBAC/quota (
ResourceQuotacan be scoped to a priority class) so only the platform can use the high ones. - Everything system-critical in kube-system uses the system classes for a reason: on a full cluster, coredns must win.
- "My pods keep getting killed for no reason" + Preempted events = somebody's priority is higher than yours.
What you can now do
- Read a FailedScheduling message node by node, and know the filter order.
- Create a PriorityClass, use it in a pod, and recognise preemption in events.