OnCallReady

Lesson 15.14 · Kubernetes: Architecture & Workloads · 28 min read

Pods: the unit, its lifecycle, and how to read its status

In plain words

A pod is like a small tent at a campsite. One or more campers (containers) share the tent: the same address on the campsite map (one IP), they can shout to each other inside (localhost), and they share the same bags (volumes). The whole tent is always pitched in one spot, never split across two fields. If the tent blows away and nobody's in charge of the campsite, nobody puts it back up.

In Kubernetes the pod is the smallest thing scheduled. It has a phase (Pending, Running, Succeeded, Failed, Unknown) and each container a state (waiting, running, terminated). The STATUS column mixes both: CrashLoopBackOff is the kubelet waiting to restart, Completed a clean exit. restartPolicy, backoff, exit codes like 137, and the SIGTERM-grace-SIGKILL shutdown are how pods live and die.

Why pods first

Every workload in Kubernetes - a web API, a database, a nightly job - ends up as pods. When something is broken, the first command anyone runs is kubectl get pods, and the answer is a STATUS word like CrashLoopBackOff or Completed. This lesson teaches what a pod is, what every column of that table means, why containers restart, where their logs are, and how they are stopped. Most of it is Docker (Ch 10-11) and systemd (Ch 2) with new names.

What you need to know already: docker run, docker logs, docker exec and exit codes (11.1, 11.5), PID 1 and signals in containers (10.28, 3.6), systemd Restart= and RestartSec= (2.10), the pod-deletion lesson (3.18), /etc/resolv.conf (8.16), the apply path and Events (15.11).

What a Pod is

A Pod is one or more containers that are scheduled together, on one node, sharing one network - one IP, one set of ports, localhost between them - and able to share files (volumes). It is the smallest thing Kubernetes schedules; never a single container.

Most pods have exactly one container, so for now "a pod is a container plus its settings" is a fine picture. (Pods with helper containers come in 15.35.)

A pod manifest (15.3: apiVersion, kind, metadata, spec):

apiVersion: v1
kind: Pod
metadata:
  name: web
  labels:
    app: web
spec:
  containers:
  - name: nginx
    image: nginx:1.27
    ports:
    - containerPort: 80       # documentation only: nothing is blocked if omitted

spec.containers is a list (the - line); each container has a name, an image and optionally the ports it listens on. Unlike docker run -p, containerPort publishes nothing - it is a label for humans and tools.

The rule: you never create bare Pods except to debug. A bare Pod has no controller (15.9). If its node dies, nothing recreates it. If it is deleted, it is gone. Everything you run for real is a pod template (a pod description) inside a Deployment, StatefulSet, DaemonSet or Job (15.16-15.24), and a controller keeps the pods coming. In this lesson you create bare pods with k run NAME --image=IMG only to see how pods behave.

Phase vs container state

A pod has a phase (a one-word summary of the whole pod) and each of its containers has a state.

Pod phase    meaning
---------    ----------------------------------------------------------------
Pending      accepted, but not all containers are running yet: unscheduled,
             pulling images, waiting on init containers or volumes
Running      bound to a node, all containers created, at least one running
             (or restarting)
Succeeded    all containers exited 0 and will not be restarted
Failed       all containers terminated, at least one non-zero, no restart
Unknown      the node stopped reporting
container state   details
---------------   --------------------------------------------------------
waiting           reason: ContainerCreating, PodInitializing, ErrImagePull,
                  ImagePullBackOff, CrashLoopBackOff, CreateContainerConfigError
running           startedAt
terminated        exitCode, reason (Completed, Error, OOMKilled), startedAt, finishedAt

The STATUS column is neither

kubectl get pods computes STATUS from both: the most interesting container reason wins, otherwise the phase.

# every STATUS at once (an illustration: these pods are not on the lab cluster)
k get pods
NAME     READY   STATUS              RESTARTS      AGE
web      1/1     Running             0             4m
bb       0/1     CrashLoopBackOff    4 (62s ago)   3m
crash    0/1     Error               2 (13s ago)   20s
waiter   0/1     Init:0/1            0             8s
new      0/1     ContainerCreating   0             2s
job-x    0/1     Completed           0             1m
gone     1/1     Terminating         0             5m

Column by column:

Later (Ch 17): readiness, liveness and startup probes - the health checks behind READY.

restartPolicy and CrashLoopBackOff

spec.restartPolicy says what the kubelet does when a container exits. It applies to all containers of the pod:

valuerestarts whenlike systemd
Always (default; the only value a Deployment allows)any exit, even 0Restart=always
OnFailurenon-zero exitRestart=on-failure
NeverneverRestart=no

When a container exits under Always, the kubelet restarts it - the first time immediately, then with exponential backoff (each wait doubles): 10s, 20s, 40s, 80s, 160s, capped at 300s. That is RestartSec= (2.10) growing on its own. The backoff resets after the container has run cleanly for 10 minutes. While waiting, the state is waiting / CrashLoopBackOff.

k run crash --image=busybox:1.36 -- sh -c '...' starts a pod crash from the small busybox image and, after --, gives it a command: print two lines, wait 2 seconds, exit with code 3:

$ k run crash --image=busybox:1.36 -- sh -c 'echo starting; echo config missing >&2; sleep 2; exit 3'
pod/crash created
$ k get pod crash
NAME    READY   STATUS             RESTARTS      AGE
crash   0/1     CrashLoopBackOff   2 (4s ago)    25s

The classic surprise: a container whose process finishes crash-loops too, because Always restarts it even on exit 0:

$ k run bb --image=busybox:1.36
$ k get pod bb
NAME   READY   STATUS             RESTARTS      AGE
bb     0/1     CrashLoopBackOff   3 (21s ago)   45s
$ k describe pod bb | sed -n '/State/,/Restart Count/p'
    State:          Waiting
      Reason:       CrashLoopBackOff
    Last State:     Terminated
      Reason:       Completed
      Exit Code:    0
    ...
    Restart Count:  3

(sed -n '/State/,/Restart Count/p' prints only the lines from "State" to "Restart Count", 7.6.) State is now; Last State is the previous run: it completed with exit code 0.

busybox's default command is sh; with no input attached it reads end-of-file and exits 0 at once. A container needs a long-running foreground process - the same rule as systemd's Type=simple (2.5). k run bb --image=busybox:1.36 -- sleep 3600 stays up.

The two logs you need

kubectl logs POD is docker logs for a pod: everything the container wrote to stdout and stderr.

$ k logs crash               # the CURRENT container instance (may be empty mid-restart)
$ k logs crash --previous    # the one that just crashed - usually what you want
starting
config missing

--previous (or -p) shows the log of the previous run. In a crash loop the current run has often printed nothing yet, so this is the one you want.

The exit code, straight from the pod's status with jsonpath (the path is status -> first container's status -> its last state -> terminated -> exitCode; 15.38 teaches the syntax):

$ k get pod crash -o jsonpath='{.status.containerStatuses[0].lastState.terminated.exitCode}{"\n"}'
3

Exit codes to recognise (the same as in 11.1): 0 finished, 1/2 the app's own error, 126 not executable, 127 command not found, 128+n killed by signal n - 137 (128+9, SIGKILL: OOMKilled, or the grace period ran out), 143 (128+15, SIGTERM: exited on shutdown).

Graceful termination

kubectl delete pod does not kill anything directly. It marks the pod for deletion - sets deletionTimestamp and a grace period, by default the pod's terminationGracePeriodSeconds: 30 - and kubectl waits for the object to disappear. On the node:

1. the pod goes Terminating; in parallel the endpoints controller removes it
   from Services (so it stops getting NEW traffic)
2. the kubelet runs the preStop hook, if any
3. the kubelet sends SIGTERM to PID 1 of each container
4. it waits up to the grace period for them to exit
5. anything still running gets SIGKILL
6. the kubelet removes the Pod object

A preStop hook is an optional command the kubelet runs inside the container just before SIGTERM (for example "sleep 5, so traffic drains first"). This is exactly the sequence from 3.18, and systemctl stop with TimeoutStopSec= (2.24).

It bites the same way too. A shell as PID 1 has no SIGTERM handler - the kernel does not apply default signal actions to PID 1 (10.28) - so this takes the full 30 seconds:

$ k run sleeper --image=busybox:1.36 -- sleep 3600
$ time k delete pod sleeper
pod "sleeper" deleted from default namespace

real	0m31.2s

(time measures how long the command took.) Its last state is Exit Code: 137: it was SIGKILLed. nginx handles SIGTERM and is gone in under a second.

Knobs on kubectl delete:

k delete pod x --grace-period=5        # shorten the wait for this delete
k delete pod x --now                   # grace period 1s
k delete pod x --force --grace-period=0   # $now: remove the object IMMEDIATELY

The last one prints a warning worth reading:

Warning: Immediate deletion does not wait for confirmation that the running resource has been terminated. The resource may continue to run on the cluster indefinitely.
pod "x" force deleted from default namespace

It deletes the API object; the container may keep running on the node until the kubelet catches up (or for ever, if the node is unreachable). For a Deployment's pod that is harmless. For a database pod with a fixed identity (a StatefulSet's db-0, 15.19) it can mean two db-0s writing to the same data - never force-delete those on a node you cannot reach.

Looking inside a running Pod

kubectl exec POD -- CMD is docker exec: run a command inside the running container. -it = interactive with a terminal, for a shell:

$ k run web --image=nginx:1.27
pod/web created
$ k exec web -- cat /etc/resolv.conf
search default.svc.cluster.local svc.cluster.local cluster.local
nameserver 10.96.0.10
options ndots:5

The pod's DNS settings: nameserver 10.96.0.10 is CoreDNS's Service IP; the search list and ndots:5 are the search-domain rules from 8.27 (short names get the cluster's domains appended).

$ k exec -it web -- sh
# hostname
web
# exit

Everything after -- is the command run inside the container; the hostname is the pod name. Old-style k exec web sh without -- is an error in current kubectl. Minimal images often lack tools (nginx has no ps; a distroless image, 10.38, has no shell at all):

# api = a distroless image (no shell inside)
k exec -it api -- sh
error: Internal error occurred: ... exec: "sh": executable file not found in $PATH: unknown

For those, kubectl debug can attach a temporary extra container with tools to a running pod (an ephemeral container); a later chapter uses it.

What you can now do:

Why it helps

You'll read kubectl get pods output all day, and every column has traps. CrashLoopBackOff isn't a reason, it's the kubelet's backoff; the reason is in logs --previous and the exit code. Exit 137 means SIGKILL: OOMKilled, or the grace period ran out. A pod taking 30 seconds to delete is usually an app ignoring SIGTERM as PID 1, the chapter 2 lesson again, and in production it means slow rollouts and dropped requests. A busybox pod crash-looping with exit code 0 teaches that containers need a long-running foreground process. These are exam staples and daily on-call material.

FAQ

What does CrashLoopBackOff actually mean?

The container keeps exiting and the kubelet is waiting before restarting it again, with exponential backoff: 10s, 20s, 40s, up to 300s, reset after 10 minutes of clean running. It isn't the crash reason. Find that with kubectl logs POD --previous and the Last State exit code in describe.

Why does my container crash-loop with exit code 0?

Under restartPolicy: Always, the default and the only option for Deployments, the kubelet restarts containers even when they exit successfully. A container whose command finishes, like busybox's default sh reading EOF, exits 0 and gets restarted forever. Containers need a long-running foreground process, the same rule as systemd's Type=simple.

Why does deleting my pod take 30 seconds?

Deletion is graceful: the kubelet sends SIGTERM to each container's PID 1, waits up to terminationGracePeriodSeconds (30 by default), then sends SIGKILL. A process running as PID 1 without a SIGTERM handler ignores the signal, since the kernel doesn't apply default signal actions to PID 1, so you wait the full grace period, and the last state shows exit code 137.

What is the difference between logs and logs --previous?

kubectl logs shows the current container instance, which in a crash loop has often just started and has little output. --previous shows the instance that just terminated, which usually contains the error or stack trace. After another restart the older instance's logs are gone, which is one reason not to restart before reading.

Should I ever create a bare pod?

Only to debug or experiment. A bare pod has no controller: if its node dies or it's deleted, nothing recreates it. Anything real runs as a pod template inside a Deployment, StatefulSet, DaemonSet or Job, so a controller keeps it alive.

In an interview Junior

A pod shows CrashLoopBackOff. How do you troubleshoot it?

CrashLoopBackOff is not the cause: it means the container keeps exiting and the kubelet is waiting before the next restart (10s, 20s, 40s ... capped at 5 minutes), as restartPolicy: Always demands.

  1. k get pod POD - RESTARTS and how long ago.
  2. k logs POD --previous - the log of the run that crashed (the current one is often empty).
  3. k describe pod POD - Last State: Terminated with the exit code and reason, plus the events.
  4. Read the exit code: 1/2 the app's own error (config, a dependency) - the logs say which; 126/127 not executable / command not found; 137 killed by SIGKILL - OOMKilled means its memory limit; 0 the process simply finished, and Always restarts even that (a container needs a long-running foreground process).
  5. Fix the cause, not the pod: the image, the command, the config or the limit in the Deployment.

Also asked: What is a Pod, and why is it the smallest unit rather than a container? · What happens step by step when a pod is deleted? · What is the difference between a pod's phase and the STATUS column?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.