OnCallReady

KubernetesSRE · 2 min read

CrashLoopBackOff: how to find out why the pod keeps restarting

CrashLoopBackOff is a symptom, not a cause. Read Last State and the exit code, get the logs of the previous container, check probes and events - a short, ordered checklist.

output
NAME                   READY   STATUS             RESTARTS      AGE
api-6c9f7b8d4c-q2xkz   0/1     CrashLoopBackOff   6 (82s ago)   9m

CrashLoopBackOff isn't an error. It's the kubelet waiting. The container exited, the kubelet restarted it, it exited again, and now each restart waits longer (10 s, 20 s, 40 s... up to 5 minutes). The cause is whatever made the container exit.

1. Read the last exit: kubectl describe pod

output
    Last State:     Terminated
      Reason:       Error
      Exit Code:    1

The exit code splits the search in two:

Exit codeUsually means
1 (or another small number)the app itself failed: config, a missing env var or secret, a dependency it can't reach
137killed by SIGKILL: Reason: OOMKilled = the memory limit, otherwise often a failed liveness probe
126 / 127the command isn't executable / not found: a wrong command: or image
0the process finished: a job image, or a command that doesn't stay in the foreground

2. Read the logs of the container that died

output
kubectl logs api-6c9f7b8d4c-q2xkz --previous

Plain kubectl logs shows the current container, which may have only just started. --previous shows the one that crashed, and it usually ends with the error.

3. Check the probes and the events

At the bottom of kubectl describe, or kubectl get events --sort-by=.lastTimestamp:

  • Liveness probe failed followed by a restart means the app may be fine but slow. A liveness probe that fires before a slow (JVM) app is up turns every start into a kill. Use a startupProbe, or a longer delay.
  • Back-off restarting failed container is the CrashLoopBackOff itself, not a new clue.

4. Then fix the real cause

  • missing config or secret: compare the pod's env and mounts with what the app expects (a key missing from a ConfigMap stops it earlier, at CreateContainerConfigError);
  • OOMKilled: see the memory limit vs the real footprint (for the JVM, heap is not the total);
  • a wrong command or entrypoint: kubectl get pod -o yaml, check command/args against the image.

The pattern is always the same: why did the last container exit? Exit code, previous logs, events. In that order. Outside Kubernetes, a Docker container that keeps restarting is the same question.

OnCallReady is free, with no ads and no tracking. RSS · All posts