The order you look in
The problem. Most Kubernetes pages are one of a handful of pod failures. If you know each one's symptom, where its evidence is and the usual fix, you stop searching the internet at 3am. This lesson is that catalogue, for pods.
What you need to know already: the debugging order (15.42), pod status and logs --previous (15.14), exit codes (3.6), OOMKilled (17.6), FailedScheduling (17.18), PVCs and StorageClasses (16.39), Jobs (15.24).
Before the catalogue, the habit. For any pod problem, in this order:
kubectl get pods -o wide # what state, which node, how many restarts
kubectl describe pod <pod> # EVENTS at the bottom - read them first
kubectl logs <pod> [-c container] # what the app said
kubectl logs <pod> --previous # what the CRASHED instance said
kubectl get events --sort-by=.lastTimestamp # everything recent in the namespace
describe Events answer 80% of pod questions in one screen. Say it early in an interview - it is the single highest-yield command in Kubernetes.
The STATUS column is a summary; the reason is one level down:
| STATUS | where the reason is |
|---|---|
| Pending | Events: FailedScheduling (scheduler), or no event at all (no scheduler) |
| ContainerCreating (long) | Events: FailedMount, FailedCreatePodSandBox (CNI) |
| ErrImagePull / ImagePullBackOff | Events: Failed to pull image ... (the registry's answer) |
| CrashLoopBackOff | Last State exit code + logs --previous |
| CreateContainerConfigError | Events: a missing ConfigMap/Secret key referenced by env |
| OOMKilled | Last State: Terminated, Reason: OOMKilled, Exit Code: 137 |
| Terminating (forever) | metadata.finalizers, or the node is gone |
CrashLoopBackOff
Symptom: RESTARTS climbing, STATUS alternating Running / Error / CrashLoopBackOff. The kubelet restarts the container with an exponential back-off (10s, 20s, 40s ... capped at 5 minutes).
# an illustration: the incidents in this chapter build each of these
kubectl get pods
NAME READY STATUS RESTARTS AGE
orders-qfpjh76znq-2wlg6 0/1 CrashLoopBackOff 3 (25s ago) 59s
Diagnosis: the exit code first, then what the previous instance printed:
# an illustration: the incidents in this chapter build each of these
kubectl describe pod orders-qfpjh76znq-2wlg6 | grep -A5 'Last State'
Last State: Terminated
Reason: Error
Exit Code: 1
kubectl logs orders-qfpjh76znq-2wlg6 --previous | tail -4
Failed to configure a DataSource: 'url' attribute is not specified and no embedded datasource could be configured.
Reason: Environment variable DB_URL is not set.
Without --previous you often get the logs of the current attempt - which may be empty because it has just started. Exit codes worth knowing:
| exit | meaning |
|---|---|
| 1 (or any app code) | the application failed: config, dependency, bug - read the logs |
| 0 | the process finished. With restartPolicy: Always (every Deployment) a finished process is restarted forever: a batch job in a Deployment, or a container whose command is sh with no script. It should be a Job, or run a server |
| 137 | 128+9, SIGKILL: OOMKilled, or killed after the grace period, or a failed liveness probe with an app that ignores SIGTERM |
| 143 | 128+15, SIGTERM: stopped - often a liveness probe restart (chapter 17) |
| 127 / 126 | command not found / not executable - usually StartError in the reason |
Fix: whatever the log says - most often config: kubectl set env deploy/orders DB_URL=..., a missing ConfigMap key, a wrong command/args. Then watch it stay up: kubectl get pods -w until RESTARTS stops increasing.
ImagePullBackOff and ErrImagePull
Symptom: ErrImagePull (a pull just failed) alternating with ImagePullBackOff (waiting to retry, with back-off).
Diagnosis: the Events contain the registry's answer, verbatim:
# an illustration: the incidents in this chapter build each of these
kubectl describe pod web-6cfbpzmhmb-cld2t | tail -4
Warning Failed 23s (x3 over 57s) kubelet Failed to pull image "registry.lab/shop/web:1.3": rpc error: code = NotFound desc = failed to pull and unpack image "registry.lab/shop/web:1.3": failed to resolve reference "registry.lab/shop/web:1.3": registry.lab/shop/web:1.3: not found
Warning Failed 23s (x3 over 57s) kubelet Error: ErrImagePull
Normal BackOff 2s (x6 over 56s) kubelet Back-off pulling image "registry.lab/shop/web:1.3"
Warning Failed 2s (x6 over 56s) kubelet Error: ImagePullBackOff
Read the tail of the first line; it tells you which of the four it is:
| message contains | cause | fix |
|---|---|---|
not found | wrong tag or repository name | fix the image reference |
pull access denied, 401 Unauthorized, authorization failed | private registry, no/wrong imagePullSecrets (the pod field naming the Secret with registry credentials) | create a docker-registry type Secret (kubectl create secret docker-registry) and reference it |
no such host, i/o timeout, dial tcp | the node cannot reach the registry: DNS, proxy, firewall | from the node: crictl pull, curl the registry |
429 Too Many Requests | Docker Hub rate limit | a pull-through cache (your own registry that fetches and keeps copies, 10.54) / mirror, or authenticate |
Fix: kubectl set image deploy/web web=registry.lab/shop/web:1.2 (or the Secret). The kubelet retries on its own - no need to delete the pod.
Pending
Symptom: Pending, no node in -o wide. The scheduler could not place it (or there is no scheduler - no events at all, 18.3).
Diagnosis: one FailedScheduling event that accounts for every node:
Warning FailedScheduling 4s (x12 over 59s) default-scheduler 0/3 nodes are available: 1 node(s) had untolerated taint {node-role.kubernetes.io/control-plane: }, 2 Insufficient cpu. preemption: 0/3 nodes are available: 1 Preemption is not helpful for scheduling, 2 No preemption victims found for incoming pod.
Read it as a sum: 3 nodes = 1 (control-plane taint) + 2 (not enough CPU). The causes, and their exact phrases:
| phrase | cause | fix |
|---|---|---|
Insufficient cpu / Insufficient memory | the requests do not fit the remaining allocatable (not actual usage!) | lower the request, add nodes, free requests elsewhere |
didn't match Pod's node affinity/selector | nodeSelector / required affinity matches no node | fix the selector, or label a node |
had untolerated taint {key: value} | the nodes that would fit are tainted | add a toleration, or schedule elsewhere |
pod has unbound immediate PersistentVolumeClaims | the PVC is Pending | look at the PVC: kubectl describe pvc, kubectl get sc |
For the PVC case, go one level further:
# an illustration: the incidents in this chapter build each of these
$ kubectl get pvc data
NAME STATUS VOLUME CAPACITY ACCESS MODES STORAGECLASS VOLUMEATTRIBUTESCLASS AGE
data Pending fast <unset> 20s
$ kubectl get sc
NAME PROVISIONER RECLAIMPOLICY VOLUMEBINDINGMODE ...
lab-disk (default) disk.csi.lab Delete WaitForFirstConsumer ...
There is no fast class. storageClassName is immutable on a PVC: delete it and recreate it with an existing class (or create the class).
Many pod fields are immutable too (resources, nodeSelector, tolerations on a bare pod): for a bare pod, edit the YAML and kubectl replace --force -f pod.yaml (delete + create). For a Deployment, edit the template - a new ReplicaSet does it.
Terminating forever
Symptom: a pod (or a whole namespace) stays Terminating long after its grace period.
# a pod stuck Terminating (the incidents build one)
$ kubectl get pod legacy -o jsonpath='{.metadata.deletionTimestamp}{" "}{.metadata.finalizers}{"\n"}'
2026-09-22T20:01:05Z ["example.com/cleanup"]
A finalizer is a key in metadata.finalizers that tells the apiserver: do not delete this object until the controller that owns this key has done its cleanup and removed it. With a deletionTimestamp set and finalizers left, the object is kept - by design. --force --grace-period=0 does not remove finalizers.
Diagnosis: who owns the finalizer, and is that controller running? A finalizer from an operator you uninstalled will wait forever.
Fix: fix or reinstall the controller if the cleanup matters. If it does not (the controller is gone for good), remove the finalizer yourself:
# an illustration: the incidents in this chapter build each of these
kubectl patch pod legacy -p '{"metadata":{"finalizers":[]}}' --type=merge
pod/legacy patched
The apiserver deletes the object as soon as the list is empty. A namespace stuck Terminating is almost always one object inside it holding a finalizer - list what is left: kubectl get all -n <ns> plus the CRDs.
The other cause: the pod's node is gone. The kubelet must confirm the containers stopped; a dead node never will (18.25). Bring the node back or delete the Node object.
OOMKilled
Symptom: restarts, and:
# an illustration: the incidents in this chapter build each of these
kubectl describe pod hog | grep -A5 'Last State'
Last State: Terminated
Reason: OOMKilled
Exit Code: 137
Diagnosis: the container exceeded its memory limit and the kernel's cgroup OOM killer killed it (5.11 and 17.6). Compare the limit with real usage: kubectl top pod (while it lives), and for a JVM the heap flags (the JVM in a container trap, 5.13).
Fix: raise the limit to what the app really needs plus headroom, or reduce what it uses (heap size, cache size, a leak). kubectl set resources deploy/x --limits=memory=512Mi. Raising the limit without understanding why is how nodes end up with MemoryPressure next week.