OnCallReady

Kubernetes · 5 min read

Pod stuck in Pending: how to read FailedScheduling (requests, taints, affinity, PVCs)

"0/3 nodes are available" is a per-node tally. Decode each reason - Insufficient cpu, untolerated taint, node affinity, unbound PVC - and fix the right one.

terminal
$ kubectl get pods -n inc-pend -o wide
NAME                       READY   STATUS    RESTARTS   AGE   IP       NODE     NOMINATED NODE   READINESS GATES
reports-z7qxh2gnh9-j52hg   0/1     Pending   0          10m   <none>   <none>   <none>           <none>

reports was deployed last night and has been Pending ever since. The first thing to check is the NODE column. <none> means the scheduler hasn't placed the pod anywhere yet. (A Pending pod that has a node is a different problem: the image is still pulling, a volume is still attaching, or a ConfigMap it mounts is missing, as in CreateContainerConfigError and FailedMount. Read that pod's events instead.)

What is actually happening

The scheduler places a pod in two passes. First it filters: every node that can't run the pod is crossed off, each for a reason. Then it scores the nodes that are left. If the filter leaves nothing, the pod stays Pending, and the scheduler writes a FailedScheduling event that lists why each node was rejected, grouped by reason. It retries whenever something changes, so once you fix the cause, the pod gets placed within seconds.

Three facts explain most of these events:

  • Requests, not usage. A pod's resources.requests is a booking. The scheduler adds up the requests of every pod on a node and compares the sum with the node's allocatable capacity. A node idling at 5% CPU still rejects a pod whose request doesn't fit what's left of its bookings.
  • Taints repel. A taint on a node (key=value:NoSchedule) keeps off every pod without a matching toleration. The control plane node has one by default, and a node that goes NotReady gets node.kubernetes.io/not-ready (one way that happens: the kubelet refusing to run with swap on).
  • Affinity attracts, and can exclude. nodeSelector and required node affinity cross off every node whose labels don't match.

Diagnosis

1. Read the event, all of it

terminal
$ kubectl describe pod -n inc-pend -l app=reports | tail -3
  Type     Reason            Age                 From               Message
  ----     ------            ----                ----               -------
  Warning  FailedScheduling  1s (x121 over 10m)  default-scheduler  0/3 nodes are available: 1 Insufficient cpu, 1 node(s) had untolerated taint {maintenance: true}, 1 node(s) had untolerated taint {node-role.kubernetes.io/control-plane: }. preemption: 0/3 nodes are available: 2 Preemption is not helpful for scheduling, 1 No preemption victims found for incoming pod.

Three nodes, three reasons. Read all of them before you fix anything. The preemption: part says the scheduler also checked whether evicting lower-priority pods would help. Here it wouldn't.

2. Translate each reason into a place to look

The message saysIt meansLook at
Insufficient cpu / Insufficient memorythe requests don't fit what's left of allocatablekubectl describe node N, Allocated resources
had untolerated taint {k: v}a taint the pod doesn't toleratekubectl describe node N | grep Taints
didn't match Pod's node affinity/selectorthe pod's nodeSelector / affinity vs the node labelskubectl get nodes --show-labels
didn't match pod anti-affinity rulesother pods' labels in the same topologythe anti-affinity term, kubectl get pods -o wide
were unschedulablethe node is cordonedkubectl get nodes (SchedulingDisabled)
pod has unbound immediate PersistentVolumeClaimsthe pod waits for its volume claimkubectl describe pvc

3. Check each node it named

terminal
$ kubectl describe node worker-1 | grep -A1 Taints
Taints:             maintenance=true:NoSchedule
Unschedulable:      false

The maintenance on worker-1 ended at 22:00, but nobody removed its taint.

terminal
$ kubectl describe node worker-2 | grep -B3 -A6 "Allocated resources"
  kube-system   coredns-d6qs9mxzfp-lmrzs   100m (5%)      0 (0%)       70Mi (1%)         170Mi (4%)      12d
  kube-system   kube-proxy-r47xw           0 (0%)         0 (0%)       0 (0%)            0 (0%)          12d
  loadtest17    soak-6mk5vt5w7p-lvl8f      1400m (70%)    0 (0%)       0 (0%)            0 (0%)          120m
Allocated resources:
  (Total limits may be over 100 percent, i.e., overcommitted.)
  Resource            Requests      Limits
  --------            --------      ------
  cpu                 1750m (87%)   0m (0%)
  memory              70Mi (1%)     170Mi (4%)

worker-2 is 87% booked. A soak test someone left running holds 1400m. reports needs 1200m, and only 250m is free.

4. Fix the cheapest real cause

The obvious reaction to "Insufficient cpu" is to add a node. That costs money and isn't needed. The taint on worker-1 is a leftover, and removing it is the step the maintenance runbook skipped:

terminal
$ kubectl taint node worker-1 maintenance=true:NoSchedule-
node/worker-1 untainted
$ kubectl get pods -n inc-pend -o wide
NAME                       READY   STATUS    RESTARTS   AGE   IP             NODE       NOMINATED NODE   READINESS GATES
reports-z7qxh2gnh9-j52hg   1/1     Running   0          10m   10.244.1.155   worker-1   <none>           <none>

Then tell the owner of the soak test that it's still running.

Two more ways to get there

A request nobody can fit. A new release asked for memory: 10Gi (someone meant 1Gi):

output
0/3 nodes are available: 1 node(s) had untolerated taint {node-role.kubernetes.io/control-plane: }, 2 Insufficient memory.
terminal
$ kubectl describe node worker-1 | grep -A8 Allocatable
Allocatable:
  cpu:                2
...
  memory:             3903128Ki

No worker offers even 4Gi. Don't just shrink it until it fits. Size it from the app's real footprint: a memory limit below that (and the limit is often set equal to the request) gets you OOMKilled pods with exit code 137 instead. Fix the request (kubectl set resources deploy/search-api -c api --requests=memory=1Gi), not the nodes. During a rolling update the old pods keep serving while the new ones sit Pending, so you have time.

A claim with no StorageClass. After a storage migration:

output
0/3 nodes are available: pod has unbound immediate PersistentVolumeClaims.

The pod is waiting for its claim, so read the claim's events:

terminal
$ kubectl describe pvc reports-data -n reports | tail -1
  Normal  FailedBinding  10s (x76 over 20m)  persistentvolume-controller  no persistent volumes available for this claim and no storage class is set

The manifests leave out storageClassName and rely on the default class, and the migration dropped the storageclass.kubernetes.io/is-default-class annotation. Mark one class as the default again. Since Kubernetes 1.28 the default is applied retroactively to existing claims that have no class, so no redeploy is needed.

How to prevent it

  • End every maintenance runbook with a check, not a hope: kubectl get nodes -o custom-columns=NAME:.metadata.name,TAINTS:.spec.taints[*].key.
  • Review requests the way you review code. A unit typo blocks a rollout just as surely as a failing test.
  • Check every new cluster for exactly one default StorageClass.
  • When a pod is Pending, read the whole FailedScheduling message. The node it mentions first is often not the cheapest one to fix.

Practise it

Chapters 15-17 run all three as incidents ("the reports job has been Pending since last night", "the new version is stuck deploying", and in the storage chapter "the new reports service is stuck Pending"), each with a different cause behind the same Pending status.

OnCallReady is free, with no ads and no tracking. RSS · All posts