Why this matters
What you need to know already: pod STATUS and READY (15.14), init containers (15.35), ConfigMap errors (15.29), Deployments and ReplicaSets (15.16), events (15.9), exec and logs from Docker (11.5), DNS checks (8.18).
At 3am "the app is down" can be twenty different things. People who guess restart things and destroy the evidence; people with a fixed order find the cause in two minutes. This lesson is that order, the same every time.
Always the same seven commands
Every "it is not working" in Kubernetes starts the same way. Do it in this order and you will find most problems in under two minutes - and say it out loud in an interview, because the interviewer is listening for exactly this.
kubectl get pods -o wide # 1. what state, which node, how many restarts
kubectl describe pod <pod> # 2. EVENTS at the bottom - read them first
kubectl logs <pod> [-c <container>] # 3. what the app says now
kubectl logs <pod> --previous # 4. what the crashed instance said
kubectl get events --sort-by=.lastTimestamp # 5. everything around it, in order
kubectl exec -it <pod> -- sh # 6. look from inside (config, DNS, files)
kubectl get deploy,rs -l app=<app> # 7. one level up: is the controller OK?
1. get -o wide: the shape of the problem
$ k get pods -o wide
NAME READY STATUS RESTARTS AGE IP NODE
api-7c9d8f6b5d-2kx8p 0/1 CrashLoopBackOff 6 (2m ago) 9m 10.244.1.23 worker-1
api-7c9d8f6b5d-n4v7q 0/1 Pending 0 9m <none> <none>
web-6b8d9c7f5d-8xk2p 0/1 Running 0 9m 10.244.2.40 worker-2
Columns: READY (ready containers / total), STATUS, RESTARTS (with how long ago), IP and NODE (-o wide adds them; <none> = not placed on a node yet). Three different problems, three different next steps:
STATUS look at
Pending, no node describe -> FailedScheduling (resources, selector, taints, PVC)
ContainerCreating describe -> FailedMount / FailedCreatePodSandBox (CNI)
ErrImagePull / ImagePullBackOff describe -> the exact pull error
CreateContainerConfigError describe -> missing ConfigMap/Secret or key
Init:N/M, Init:CrashLoopBackOff logs -c <init container>
CrashLoopBackOff / Error logs --previous, Last State exit code
Running but 0/1 describe -> Readiness probe failed
Terminating for a long time finalizers, the node, a process ignoring SIGTERM
(no pods at all) describe rs -> FailedCreate
Words in that table you meet properly later: a taint is a "keep off" mark on a node (the control plane has one, Ch 17); a PVC is a request for storage (Ch 16); a readiness probe is a health check that decides whether a pod gets traffic (Ch 17); a finalizer is a note on an object that blocks its deletion until some controller has cleaned up (Ch 16). FailedScheduling = the scheduler found no node; FailedCreatePodSandBox = the pod's network could not be set up.
2. describe: read it bottom-up
The Events table at the bottom is the single highest-yield output in Kubernetes. Then, going up: Last State (exit code and reason of the previous crash), State, Ready, Restart Count, the Conditions table, and the env and mounts the container actually got.
3 and 4. logs, and which instance
kubectl logs shows the current container instance. In a crash loop that is usually a container that has just started (or has not started), so its log is short or empty. --previous is the one that died - the one with the stack trace.
k logs deploy/api # picks one pod of the Deployment
k logs -l app=api --prefix # every pod, prefixed with pod/container
k logs api-xyz -c migrate # a specific container (init containers too)
k logs api-xyz --tail=50 -f # follow
k logs api-xyz --since=10m --timestamps
5. events for context
k get events --sort-by=.lastTimestamp (or k events) shows what happened around the pod: a node that went NotReady, a ReplicaSet that could not create pods, a scale-down. -A --field-selector=type=Warning across the whole cluster is the fastest "what is wrong right now" query there is. Remember they expire after an hour.
6. exec: check the assumptions
k exec -it api-xyz -- sh
env | sort # did the config arrive?
cat /etc/app/application.yaml # is the mounted file what you think?
cat /etc/resolv.conf # DNS search and nameserver
nslookup db # does the dependency resolve?
wget -qO- http://db:8080/health # is it reachable from here?
If the image has no shell (distroless, 10.38), kubectl debug attaches a temporary extra container with tools (Ch 18).
7. one level up
A pod problem is sometimes a controller problem: the Deployment is paused, the rollout hit its progress deadline, the ReplicaSet cannot create pods, the node the pods are on is NotReady. get deploy,rs and describe deploy (Conditions, and OldReplicaSets/NewReplicaSet) answer that.
The habit
Do not guess and do not restart first. Restarting destroys the evidence (the previous logs are gone after another restart, events expire). Read, then act.
What you can now do:
- Run the seven commands in order and know what each one answers.
- Map a pod STATUS to the next command to run.
- Go one level up to the ReplicaSet and Deployment when no pods appear.