OnCallReady

Lesson 18.28 · Kubernetes: Cluster Operations & Troubleshooting · 19 min read

The failure catalogue I: pods that will not run

In plain words

Think of a doctor with a list of common illnesses. Each one has a telltale symptom and a first question to ask. A fever and a rash? Ask about measles. A cough at night? Ask about the dog. The doctor doesn't guess; they look at the chart first, then ask the right question.

This lesson is that list for pods. The chart is kubectl describe pod, with its Events and Last State. CrashLoopBackOff: read the exit code and logs --previous. ImagePullBackOff: read the registry's exact answer. Pending: read FailedScheduling as a sum over nodes. Terminating forever: look for a finalizer or a dead node. OOMKilled: compare the memory limit with real usage.

The order you look in

The problem. Most Kubernetes pages are one of a handful of pod failures. If you know each one's symptom, where its evidence is and the usual fix, you stop searching the internet at 3am. This lesson is that catalogue, for pods.

What you need to know already: the debugging order (15.42), pod status and logs --previous (15.14), exit codes (3.6), OOMKilled (17.6), FailedScheduling (17.18), PVCs and StorageClasses (16.39), Jobs (15.24).

Before the catalogue, the habit. For any pod problem, in this order:

kubectl get pods -o wide                      # what state, which node, how many restarts
kubectl describe pod <pod>                    # EVENTS at the bottom - read them first
kubectl logs <pod> [-c container]             # what the app said
kubectl logs <pod> --previous                 # what the CRASHED instance said
kubectl get events --sort-by=.lastTimestamp   # everything recent in the namespace

describe Events answer 80% of pod questions in one screen. Say it early in an interview - it is the single highest-yield command in Kubernetes.

The STATUS column is a summary; the reason is one level down:

STATUSwhere the reason is
PendingEvents: FailedScheduling (scheduler), or no event at all (no scheduler)
ContainerCreating (long)Events: FailedMount, FailedCreatePodSandBox (CNI)
ErrImagePull / ImagePullBackOffEvents: Failed to pull image ... (the registry's answer)
CrashLoopBackOffLast State exit code + logs --previous
CreateContainerConfigErrorEvents: a missing ConfigMap/Secret key referenced by env
OOMKilledLast State: Terminated, Reason: OOMKilled, Exit Code: 137
Terminating (forever)metadata.finalizers, or the node is gone

CrashLoopBackOff

Symptom: RESTARTS climbing, STATUS alternating Running / Error / CrashLoopBackOff. The kubelet restarts the container with an exponential back-off (10s, 20s, 40s ... capped at 5 minutes).

# an illustration: the incidents in this chapter build each of these
kubectl get pods
NAME                      READY   STATUS             RESTARTS      AGE
orders-qfpjh76znq-2wlg6   0/1     CrashLoopBackOff   3 (25s ago)   59s

Diagnosis: the exit code first, then what the previous instance printed:

# an illustration: the incidents in this chapter build each of these
kubectl describe pod orders-qfpjh76znq-2wlg6 | grep -A5 'Last State'
    Last State:     Terminated
      Reason:       Error
      Exit Code:    1
kubectl logs orders-qfpjh76znq-2wlg6 --previous | tail -4
Failed to configure a DataSource: 'url' attribute is not specified and no embedded datasource could be configured.

Reason: Environment variable DB_URL is not set.

Without --previous you often get the logs of the current attempt - which may be empty because it has just started. Exit codes worth knowing:

exitmeaning
1 (or any app code)the application failed: config, dependency, bug - read the logs
0the process finished. With restartPolicy: Always (every Deployment) a finished process is restarted forever: a batch job in a Deployment, or a container whose command is sh with no script. It should be a Job, or run a server
137128+9, SIGKILL: OOMKilled, or killed after the grace period, or a failed liveness probe with an app that ignores SIGTERM
143128+15, SIGTERM: stopped - often a liveness probe restart (chapter 17)
127 / 126command not found / not executable - usually StartError in the reason

Fix: whatever the log says - most often config: kubectl set env deploy/orders DB_URL=..., a missing ConfigMap key, a wrong command/args. Then watch it stay up: kubectl get pods -w until RESTARTS stops increasing.

ImagePullBackOff and ErrImagePull

Symptom: ErrImagePull (a pull just failed) alternating with ImagePullBackOff (waiting to retry, with back-off).

Diagnosis: the Events contain the registry's answer, verbatim:

# an illustration: the incidents in this chapter build each of these
kubectl describe pod web-6cfbpzmhmb-cld2t | tail -4
  Warning  Failed     23s (x3 over 57s)  kubelet  Failed to pull image "registry.lab/shop/web:1.3": rpc error: code = NotFound desc = failed to pull and unpack image "registry.lab/shop/web:1.3": failed to resolve reference "registry.lab/shop/web:1.3": registry.lab/shop/web:1.3: not found
  Warning  Failed     23s (x3 over 57s)  kubelet  Error: ErrImagePull
  Normal   BackOff    2s (x6 over 56s)   kubelet  Back-off pulling image "registry.lab/shop/web:1.3"
  Warning  Failed     2s (x6 over 56s)   kubelet  Error: ImagePullBackOff

Read the tail of the first line; it tells you which of the four it is:

message containscausefix
not foundwrong tag or repository namefix the image reference
pull access denied, 401 Unauthorized, authorization failedprivate registry, no/wrong imagePullSecrets (the pod field naming the Secret with registry credentials)create a docker-registry type Secret (kubectl create secret docker-registry) and reference it
no such host, i/o timeout, dial tcpthe node cannot reach the registry: DNS, proxy, firewallfrom the node: crictl pull, curl the registry
429 Too Many RequestsDocker Hub rate limita pull-through cache (your own registry that fetches and keeps copies, 10.54) / mirror, or authenticate

Fix: kubectl set image deploy/web web=registry.lab/shop/web:1.2 (or the Secret). The kubelet retries on its own - no need to delete the pod.

Pending

Symptom: Pending, no node in -o wide. The scheduler could not place it (or there is no scheduler - no events at all, 18.3).

Diagnosis: one FailedScheduling event that accounts for every node:

  Warning  FailedScheduling  4s (x12 over 59s)  default-scheduler  0/3 nodes are available: 1 node(s) had untolerated taint {node-role.kubernetes.io/control-plane: }, 2 Insufficient cpu. preemption: 0/3 nodes are available: 1 Preemption is not helpful for scheduling, 2 No preemption victims found for incoming pod.

Read it as a sum: 3 nodes = 1 (control-plane taint) + 2 (not enough CPU). The causes, and their exact phrases:

phrasecausefix
Insufficient cpu / Insufficient memorythe requests do not fit the remaining allocatable (not actual usage!)lower the request, add nodes, free requests elsewhere
didn't match Pod's node affinity/selectornodeSelector / required affinity matches no nodefix the selector, or label a node
had untolerated taint {key: value}the nodes that would fit are taintedadd a toleration, or schedule elsewhere
pod has unbound immediate PersistentVolumeClaimsthe PVC is Pendinglook at the PVC: kubectl describe pvc, kubectl get sc

For the PVC case, go one level further:

# an illustration: the incidents in this chapter build each of these
$ kubectl get pvc data
NAME   STATUS    VOLUME   CAPACITY   ACCESS MODES   STORAGECLASS   VOLUMEATTRIBUTESCLASS   AGE
data   Pending                                      fast           <unset>                 20s
$ kubectl get sc
NAME                 PROVISIONER             RECLAIMPOLICY   VOLUMEBINDINGMODE      ...
lab-disk (default)   disk.csi.lab            Delete          WaitForFirstConsumer   ...

There is no fast class. storageClassName is immutable on a PVC: delete it and recreate it with an existing class (or create the class).

Many pod fields are immutable too (resources, nodeSelector, tolerations on a bare pod): for a bare pod, edit the YAML and kubectl replace --force -f pod.yaml (delete + create). For a Deployment, edit the template - a new ReplicaSet does it.

Terminating forever

Symptom: a pod (or a whole namespace) stays Terminating long after its grace period.

# a pod stuck Terminating (the incidents build one)
$ kubectl get pod legacy -o jsonpath='{.metadata.deletionTimestamp}{"  "}{.metadata.finalizers}{"\n"}'
2026-09-22T20:01:05Z  ["example.com/cleanup"]

A finalizer is a key in metadata.finalizers that tells the apiserver: do not delete this object until the controller that owns this key has done its cleanup and removed it. With a deletionTimestamp set and finalizers left, the object is kept - by design. --force --grace-period=0 does not remove finalizers.

Diagnosis: who owns the finalizer, and is that controller running? A finalizer from an operator you uninstalled will wait forever.

Fix: fix or reinstall the controller if the cleanup matters. If it does not (the controller is gone for good), remove the finalizer yourself:

# an illustration: the incidents in this chapter build each of these
kubectl patch pod legacy -p '{"metadata":{"finalizers":[]}}' --type=merge
pod/legacy patched

The apiserver deletes the object as soon as the list is empty. A namespace stuck Terminating is almost always one object inside it holding a finalizer - list what is left: kubectl get all -n <ns> plus the CRDs.

The other cause: the pod's node is gone. The kubelet must confirm the containers stopped; a dead node never will (18.25). Bring the node back or delete the Node object.

OOMKilled

Symptom: restarts, and:

# an illustration: the incidents in this chapter build each of these
kubectl describe pod hog | grep -A5 'Last State'
    Last State:     Terminated
      Reason:       OOMKilled
      Exit Code:    137

Diagnosis: the container exceeded its memory limit and the kernel's cgroup OOM killer killed it (5.11 and 17.6). Compare the limit with real usage: kubectl top pod (while it lives), and for a JVM the heap flags (the JVM in a container trap, 5.13).

Fix: raise the limit to what the app really needs plus headroom, or reduce what it uses (heap size, cache size, a leak). kubectl set resources deploy/x --limits=memory=512Mi. Raising the limit without understanding why is how nodes end up with MemoryPressure next week.

Why it helps

These five are the bulk of the tickets any platform team gets, and each has a two-command diagnosis. Being fast here is what makes you trusted: "CrashLoopBackOff, exit 1, logs --previous says DB_URL isn't set" in thirty seconds instead of a round of guessing. Exit codes (0, 1, 137, 143, 127) become a vocabulary: 0 in a Deployment means a batch job in the wrong kind of controller; 137 means SIGKILL.

On the admin exam, the troubleshooting domain is the largest, and these are its bread and butter. In interviews, "a pod is in CrashLoopBackOff, what do you do?" is asked constantly, and "describe Events first, then logs --previous" is the answer interviewers want to hear.

Commands in this lesson

kubectl

FAQ

Why are the logs empty for my crash-looping pod?

You're probably reading the current attempt, which has just started or not started yet. Use kubectl logs <pod> --previous to read what the crashed instance printed before it died, and add -c <container> for multi-container pods. Also read Last State in describe for the exit code and reason.

My container exits with code 0 but the pod is in CrashLoopBackOff. How?

Deployments use restartPolicy: Always, so a process that finishes successfully is restarted forever, and repeated restarts trigger back-off. Either it's genuinely a one-off task, which should be a Job, or the container's command ends immediately (a shell with no script, a wrong entrypoint) when it should run a server.

Do I need to delete the pod after fixing an image pull problem?

No. The kubelet keeps retrying with back-off (that's what ImagePullBackOff means). Fix the reference with kubectl set image (which rolls out new pods) or create the missing imagePullSecrets Secret, and the next retry succeeds. For a bare pod with a wrong image, you can edit the image field in place; it's one of the few mutable fields.

Why doesn't --force --grace-period=0 delete my Terminating pod?

If the pod has finalizers, the API server keeps the object until every finalizer is removed; force only skips the graceful wait. Check kubectl get pod -o jsonpath='{.metadata.finalizers}'. If the controller owning the finalizer is gone for good, remove it with a merge patch setting finalizers to an empty list. If the node is dead, that's the other cause.

How do I change a field on a bare pod that says it's immutable?

Most pod spec fields (resources, nodeSelector, tolerations, volumes) can't be changed on a running pod. Export it, edit the YAML and run kubectl replace --force -f pod.yaml, which deletes and recreates it. For a Deployment, edit the pod template instead; the new ReplicaSet does the replacement.

In an interview Junior

A pod is stuck in ImagePullBackOff. How do you diagnose it?

ErrImagePull is a failed pull; ImagePullBackOff is the kubelet waiting to retry. kubectl describe pod POD - the Events contain the registry's answer verbatim. Read the end of it:

The kubelet keeps retrying with back-off, so once the cause is fixed there is no need to delete the pod.

Also asked: A pod is in CrashLoopBackOff. What do you do? · A namespace has been stuck in Terminating for a day. How do you resolve it? · What is a finalizer?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.