OnCallReady

Lesson 15.42 · Kubernetes: Architecture & Workloads · 11 min read

The debugging order

In plain words

When a doctor sees a patient who says "I don't feel well", they don't guess or start with surgery. They follow the same routine every time: look at the patient, ask what happened, check the notes from the last visit, look at what's going on around them, examine them closely, and check whether the problem is actually somewhere else. And they never give medicine before examining, because that would hide the symptoms.

Debugging in Kubernetes has the same routine: kubectl get pods -o wide (state, node, restarts), describe pod (Events at the bottom first), logs, logs --previous (the crashed instance), get events --sort-by=.lastTimestamp, exec inside to check assumptions, and one level up (get deploy,rs). Read before you act: restarting destroys the evidence.

Why this matters

What you need to know already: pod STATUS and READY (15.14), init containers (15.35), ConfigMap errors (15.29), Deployments and ReplicaSets (15.16), events (15.9), exec and logs from Docker (11.5), DNS checks (8.18).

At 3am "the app is down" can be twenty different things. People who guess restart things and destroy the evidence; people with a fixed order find the cause in two minutes. This lesson is that order, the same every time.

Always the same seven commands

Every "it is not working" in Kubernetes starts the same way. Do it in this order and you will find most problems in under two minutes - and say it out loud in an interview, because the interviewer is listening for exactly this.

kubectl get pods -o wide                       # 1. what state, which node, how many restarts
kubectl describe pod <pod>                     # 2. EVENTS at the bottom - read them first
kubectl logs <pod> [-c <container>]            # 3. what the app says now
kubectl logs <pod> --previous                  # 4. what the crashed instance said
kubectl get events --sort-by=.lastTimestamp    # 5. everything around it, in order
kubectl exec -it <pod> -- sh                   # 6. look from inside (config, DNS, files)
kubectl get deploy,rs -l app=<app>             # 7. one level up: is the controller OK?

1. get -o wide: the shape of the problem

$ k get pods -o wide
NAME                   READY   STATUS             RESTARTS      AGE   IP             NODE
api-7c9d8f6b5d-2kx8p   0/1     CrashLoopBackOff   6 (2m ago)    9m    10.244.1.23    worker-1
api-7c9d8f6b5d-n4v7q   0/1     Pending            0             9m    <none>         <none>
web-6b8d9c7f5d-8xk2p   0/1     Running            0             9m    10.244.2.40    worker-2

Columns: READY (ready containers / total), STATUS, RESTARTS (with how long ago), IP and NODE (-o wide adds them; <none> = not placed on a node yet). Three different problems, three different next steps:

STATUS                          look at
Pending, no node                describe -> FailedScheduling (resources, selector, taints, PVC)
ContainerCreating               describe -> FailedMount / FailedCreatePodSandBox (CNI)
ErrImagePull / ImagePullBackOff describe -> the exact pull error
CreateContainerConfigError      describe -> missing ConfigMap/Secret or key
Init:N/M, Init:CrashLoopBackOff logs -c <init container>
CrashLoopBackOff / Error        logs --previous, Last State exit code
Running but 0/1                 describe -> Readiness probe failed
Terminating for a long time     finalizers, the node, a process ignoring SIGTERM
(no pods at all)                describe rs -> FailedCreate

Words in that table you meet properly later: a taint is a "keep off" mark on a node (the control plane has one, Ch 17); a PVC is a request for storage (Ch 16); a readiness probe is a health check that decides whether a pod gets traffic (Ch 17); a finalizer is a note on an object that blocks its deletion until some controller has cleaned up (Ch 16). FailedScheduling = the scheduler found no node; FailedCreatePodSandBox = the pod's network could not be set up.

2. describe: read it bottom-up

The Events table at the bottom is the single highest-yield output in Kubernetes. Then, going up: Last State (exit code and reason of the previous crash), State, Ready, Restart Count, the Conditions table, and the env and mounts the container actually got.

3 and 4. logs, and which instance

kubectl logs shows the current container instance. In a crash loop that is usually a container that has just started (or has not started), so its log is short or empty. --previous is the one that died - the one with the stack trace.

k logs deploy/api                 # picks one pod of the Deployment
k logs -l app=api --prefix        # every pod, prefixed with pod/container
k logs api-xyz -c migrate         # a specific container (init containers too)
k logs api-xyz --tail=50 -f       # follow
k logs api-xyz --since=10m --timestamps

5. events for context

k get events --sort-by=.lastTimestamp (or k events) shows what happened around the pod: a node that went NotReady, a ReplicaSet that could not create pods, a scale-down. -A --field-selector=type=Warning across the whole cluster is the fastest "what is wrong right now" query there is. Remember they expire after an hour.

6. exec: check the assumptions

k exec -it api-xyz -- sh
env | sort                     # did the config arrive?
cat /etc/app/application.yaml  # is the mounted file what you think?
cat /etc/resolv.conf           # DNS search and nameserver
nslookup db                    # does the dependency resolve?
wget -qO- http://db:8080/health   # is it reachable from here?

If the image has no shell (distroless, 10.38), kubectl debug attaches a temporary extra container with tools (Ch 18).

7. one level up

A pod problem is sometimes a controller problem: the Deployment is paused, the rollout hit its progress deadline, the ReplicaSet cannot create pods, the node the pods are on is NotReady. get deploy,rs and describe deploy (Conditions, and OldReplicaSets/NewReplicaSet) answer that.

The habit

Do not guess and do not restart first. Restarting destroys the evidence (the previous logs are gone after another restart, events expire). Read, then act.

What you can now do:

Why it helps

This is what an interviewer is listening for when they ask "a pod isn't working, what do you do?", and what you'll do on every call in your first months on the platform team. The STATUS column routes you: Pending means scheduling, ContainerCreating means mounts or CNI, ImagePullBackOff means the image, CreateContainerConfigError means missing config, CrashLoopBackOff means the app itself, Running 0/1 means readiness. The discipline of not restarting first matters in real incidents: after another restart, the previous logs are gone, and events expire after an hour. In the exam, troubleshooting is a large share of the score.

FAQ

Why read the events first?

Because every component that acted on the pod recorded what it did or why it failed, and the last events tell you which hop stopped: scheduler, kubelet mount, image pull, container start, probes. It's the highest-yield output in Kubernetes. Read describe bottom-up: Events, then Last State and exit code, then State, Ready and conditions.

The pod is Running but READY 0/1. What does that mean?

The container is running but its readiness probe is failing, so the pod is excluded from Service endpoints and receives no traffic. describe shows "Readiness probe failed" with the reason, like a connection refused or an HTTP 500. Check the probe's port and path against what the app actually serves, and whether the app is waiting for a dependency.

What if the image has no shell for kubectl exec?

Minimal and distroless images often have no sh, so kubectl exec -it POD -- sh fails with "executable file not found". Use kubectl debug to attach an ephemeral container with tools that shares the pod's namespaces (chapter 18), or run a separate debug pod in the same namespace to test DNS and connectivity.

Why shouldn't I just restart the pod?

Restarting destroys evidence. The crashed instance's logs, available through --previous, are gone after another restart, events expire after an hour, and the state that caused the problem may not recur immediately. Capture first: describe, logs, previous logs, events. Then act. If you must restart to restore service, save the output before you do.

What if there are no pods at all?

Go one level up. kubectl get deploy,rs -l app=NAME and kubectl describe rs show whether the ReplicaSet could create pods; FailedCreate events reveal quota limits, missing ServiceAccounts or admission rejections. describe deploy shows whether the Deployment is paused or has exceeded its progress deadline.

In an interview Junior

A pod is not working. What are your first steps?

The same order every time, reading before acting:

  1. k get pods -o wide - STATUS, READY, RESTARTS and the node. Pending = scheduling; ContainerCreating = sandbox or mounts; ImagePullBackOff = image; CrashLoopBackOff = the app exits; Running but 0/1 = readiness.
  2. k describe pod POD - the Events at the bottom first, then Last State (exit code, reason).
  3. k logs POD and 4. k logs POD --previous - the crashed run usually has the error.
  4. k get events --sort-by=.lastTimestamp - what happened around it (they expire after an hour).
  5. k exec -it POD -- sh - check assumptions: env, files, DNS, the port.
  6. One level up - k get deploy,rs and describe deploy: a paused or stuck rollout, a ReplicaSet that cannot create pods.

Do not restart first: it destroys the evidence (previous logs, events).

Also asked: Map the common pod statuses to their likely causes. · What does FailedScheduling tell you? · How do you debug a container image that has no shell?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.