Why pods first
Every workload in Kubernetes - a web API, a database, a nightly job - ends up as pods. When something is broken, the first command anyone runs is kubectl get pods, and the answer is a STATUS word like CrashLoopBackOff or Completed. This lesson teaches what a pod is, what every column of that table means, why containers restart, where their logs are, and how they are stopped. Most of it is Docker (Ch 10-11) and systemd (Ch 2) with new names.
What you need to know already: docker run, docker logs, docker exec and exit codes (11.1, 11.5), PID 1 and signals in containers (10.28, 3.6), systemd Restart= and RestartSec= (2.10), the pod-deletion lesson (3.18), /etc/resolv.conf (8.16), the apply path and Events (15.11).
What a Pod is
A Pod is one or more containers that are scheduled together, on one node, sharing one network - one IP, one set of ports, localhost between them - and able to share files (volumes). It is the smallest thing Kubernetes schedules; never a single container.
Most pods have exactly one container, so for now "a pod is a container plus its settings" is a fine picture. (Pods with helper containers come in 15.35.)
A pod manifest (15.3: apiVersion, kind, metadata, spec):
apiVersion: v1
kind: Pod
metadata:
name: web
labels:
app: web
spec:
containers:
- name: nginx
image: nginx:1.27
ports:
- containerPort: 80 # documentation only: nothing is blocked if omitted
spec.containers is a list (the - line); each container has a name, an image and optionally the ports it listens on. Unlike docker run -p, containerPort publishes nothing - it is a label for humans and tools.
The rule: you never create bare Pods except to debug. A bare Pod has no controller (15.9). If its node dies, nothing recreates it. If it is deleted, it is gone. Everything you run for real is a pod template (a pod description) inside a Deployment, StatefulSet, DaemonSet or Job (15.16-15.24), and a controller keeps the pods coming. In this lesson you create bare pods with k run NAME --image=IMG only to see how pods behave.
Phase vs container state
A pod has a phase (a one-word summary of the whole pod) and each of its containers has a state.
Pod phase meaning
--------- ----------------------------------------------------------------
Pending accepted, but not all containers are running yet: unscheduled,
pulling images, waiting on init containers or volumes
Running bound to a node, all containers created, at least one running
(or restarting)
Succeeded all containers exited 0 and will not be restarted
Failed all containers terminated, at least one non-zero, no restart
Unknown the node stopped reporting
container state details
--------------- --------------------------------------------------------
waiting reason: ContainerCreating, PodInitializing, ErrImagePull,
ImagePullBackOff, CrashLoopBackOff, CreateContainerConfigError
running startedAt
terminated exitCode, reason (Completed, Error, OOMKilled), startedAt, finishedAt
- waiting - not running yet, with a reason: still being created, the image cannot be pulled (ErrImagePull, then ImagePullBackOff while the kubelet waits to retry), waiting between restarts (CrashLoopBackOff, below), or its settings point at something missing (CreateContainerConfigError, 15.29).
- running - the process is up.
- terminated - it exited: the exit code and a reason. OOMKilled means the kernel killed it for going over its memory limit (the cgroup OOM from 5.11).
The STATUS column is neither
kubectl get pods computes STATUS from both: the most interesting container reason wins, otherwise the phase.
# every STATUS at once (an illustration: these pods are not on the lab cluster)
k get pods
NAME READY STATUS RESTARTS AGE
web 1/1 Running 0 4m
bb 0/1 CrashLoopBackOff 4 (62s ago) 3m
crash 0/1 Error 2 (13s ago) 20s
waiter 0/1 Init:0/1 0 8s
new 0/1 ContainerCreating 0 2s
job-x 0/1 Completed 0 1m
gone 1/1 Terminating 0 5m
Column by column:
- NAME - the pod's name.
- READY
1/1- ready containers / containers. Not "running": a container can run but fail its readiness check (a health check the kubelet runs), and then it counts as 0. Only ready pods receive traffic. - STATUS - the summary word: ·
Running- all good (as far as the process is concerned). ·CrashLoopBackOff- the container keeps exiting and the kubelet is waiting before the next restart. It is not a crash reason; it is the kubelet's pause. ·Error/Completed- the container's last exit: non-zero / zero. You see it between restarts, or for good when the pod is done. ·Init:0/1- 0 of 1 setup containers (init containers, 15.35) done. ·ContainerCreating- the kubelet is setting it up (network, image pull). ·Terminating- someone deleted it; graceful shutdown in progress. - RESTARTS
4 (62s ago)- how many times its containers were restarted in place, and how long ago the last one ended. - AGE - how long ago the pod object was created.
Later (Ch 17): readiness, liveness and startup probes - the health checks behind READY.
restartPolicy and CrashLoopBackOff
spec.restartPolicy says what the kubelet does when a container exits. It applies to all containers of the pod:
| value | restarts when | like systemd |
|---|---|---|
Always (default; the only value a Deployment allows) | any exit, even 0 | Restart=always |
OnFailure | non-zero exit | Restart=on-failure |
Never | never | Restart=no |
When a container exits under Always, the kubelet restarts it - the first time immediately, then with exponential backoff (each wait doubles): 10s, 20s, 40s, 80s, 160s, capped at 300s. That is RestartSec= (2.10) growing on its own. The backoff resets after the container has run cleanly for 10 minutes. While waiting, the state is waiting / CrashLoopBackOff.
k run crash --image=busybox:1.36 -- sh -c '...' starts a pod crash from the small busybox image and, after --, gives it a command: print two lines, wait 2 seconds, exit with code 3:
$ k run crash --image=busybox:1.36 -- sh -c 'echo starting; echo config missing >&2; sleep 2; exit 3'
pod/crash created
$ k get pod crash
NAME READY STATUS RESTARTS AGE
crash 0/1 CrashLoopBackOff 2 (4s ago) 25s
The classic surprise: a container whose process finishes crash-loops too, because Always restarts it even on exit 0:
$ k run bb --image=busybox:1.36
$ k get pod bb
NAME READY STATUS RESTARTS AGE
bb 0/1 CrashLoopBackOff 3 (21s ago) 45s
$ k describe pod bb | sed -n '/State/,/Restart Count/p'
State: Waiting
Reason: CrashLoopBackOff
Last State: Terminated
Reason: Completed
Exit Code: 0
...
Restart Count: 3
(sed -n '/State/,/Restart Count/p' prints only the lines from "State" to "Restart Count", 7.6.) State is now; Last State is the previous run: it completed with exit code 0.
busybox's default command is sh; with no input attached it reads end-of-file and exits 0 at once. A container needs a long-running foreground process - the same rule as systemd's Type=simple (2.5). k run bb --image=busybox:1.36 -- sleep 3600 stays up.
The two logs you need
kubectl logs POD is docker logs for a pod: everything the container wrote to stdout and stderr.
$ k logs crash # the CURRENT container instance (may be empty mid-restart)
$ k logs crash --previous # the one that just crashed - usually what you want
starting
config missing
--previous (or -p) shows the log of the previous run. In a crash loop the current run has often printed nothing yet, so this is the one you want.
The exit code, straight from the pod's status with jsonpath (the path is status -> first container's status -> its last state -> terminated -> exitCode; 15.38 teaches the syntax):
$ k get pod crash -o jsonpath='{.status.containerStatuses[0].lastState.terminated.exitCode}{"\n"}'
3
Exit codes to recognise (the same as in 11.1): 0 finished, 1/2 the app's own error, 126 not executable, 127 command not found, 128+n killed by signal n - 137 (128+9, SIGKILL: OOMKilled, or the grace period ran out), 143 (128+15, SIGTERM: exited on shutdown).
Graceful termination
kubectl delete pod does not kill anything directly. It marks the pod for deletion - sets deletionTimestamp and a grace period, by default the pod's terminationGracePeriodSeconds: 30 - and kubectl waits for the object to disappear. On the node:
1. the pod goes Terminating; in parallel the endpoints controller removes it
from Services (so it stops getting NEW traffic)
2. the kubelet runs the preStop hook, if any
3. the kubelet sends SIGTERM to PID 1 of each container
4. it waits up to the grace period for them to exit
5. anything still running gets SIGKILL
6. the kubelet removes the Pod object
A preStop hook is an optional command the kubelet runs inside the container just before SIGTERM (for example "sleep 5, so traffic drains first"). This is exactly the sequence from 3.18, and systemctl stop with TimeoutStopSec= (2.24).
It bites the same way too. A shell as PID 1 has no SIGTERM handler - the kernel does not apply default signal actions to PID 1 (10.28) - so this takes the full 30 seconds:
$ k run sleeper --image=busybox:1.36 -- sleep 3600
$ time k delete pod sleeper
pod "sleeper" deleted from default namespace
real 0m31.2s
(time measures how long the command took.) Its last state is Exit Code: 137: it was SIGKILLed. nginx handles SIGTERM and is gone in under a second.
Knobs on kubectl delete:
k delete pod x --grace-period=5 # shorten the wait for this delete
k delete pod x --now # grace period 1s
k delete pod x --force --grace-period=0 # $now: remove the object IMMEDIATELY
The last one prints a warning worth reading:
Warning: Immediate deletion does not wait for confirmation that the running resource has been terminated. The resource may continue to run on the cluster indefinitely.
pod "x" force deleted from default namespace
It deletes the API object; the container may keep running on the node until the kubelet catches up (or for ever, if the node is unreachable). For a Deployment's pod that is harmless. For a database pod with a fixed identity (a StatefulSet's db-0, 15.19) it can mean two db-0s writing to the same data - never force-delete those on a node you cannot reach.
Looking inside a running Pod
kubectl exec POD -- CMD is docker exec: run a command inside the running container. -it = interactive with a terminal, for a shell:
$ k run web --image=nginx:1.27
pod/web created
$ k exec web -- cat /etc/resolv.conf
search default.svc.cluster.local svc.cluster.local cluster.local
nameserver 10.96.0.10
options ndots:5
The pod's DNS settings: nameserver 10.96.0.10 is CoreDNS's Service IP; the search list and ndots:5 are the search-domain rules from 8.27 (short names get the cluster's domains appended).
$ k exec -it web -- sh
# hostname
web
# exit
Everything after -- is the command run inside the container; the hostname is the pod name. Old-style k exec web sh without -- is an error in current kubectl. Minimal images often lack tools (nginx has no ps; a distroless image, 10.38, has no shell at all):
# api = a distroless image (no shell inside)
k exec -it api -- sh
error: Internal error occurred: ... exec: "sh": executable file not found in $PATH: unknown
For those, kubectl debug can attach a temporary extra container with tools to a running pod (an ephemeral container); a later chapter uses it.
What you can now do:
- read every column of
kubectl get pods, and phase vs container state - explain CrashLoopBackOff and restartPolicy in systemd terms
- get the logs of the crashed run, the exit code, and a shell inside a pod
- predict how long a delete takes and why