$ kubectl set image deploy/storefront -n front web=nginx:1.28
deployment.apps/storefront image updated
$ kubectl get pods -n front
NAME READY STATUS RESTARTS AGE
storefront-qzpq58ml9h-ffnb4 1/1 Terminating 0 10m
storefront-qzpq58ml9h-gmwmb 1/1 Terminating 0 10m
storefront-qzpq58ml9h-ww5bk 1/1 Terminating 0 10m
$ kubectl get pods -n front
NAME READY STATUS RESTARTS AGE
storefront-6w59z2cf6l-cs99c 0/1 ContainerCreating 0 1s
storefront-6w59z2cf6l-qpjhl 0/1 ContainerCreating 0 1s
storefront-6w59z2cf6l-w5d96 0/1 ContainerCreating 0 1sEvery deploy of storefront pages the uptime monitor: 20 to 40 seconds of 503s. It's nginx, and it starts in two seconds. So why is there any downtime at all?
Look at the two listings. First every old pod is Terminating, then every new pod is still being created. For a few seconds in between, no pod at all can serve.
What is actually happening
A deploy can lose requests at three different points. Check them in this order.
- The strategy removes old pods before new ones are ready.
strategy: Recreatekills every old pod, waits until they're all gone (up to their grace period), and only then creates the new ones. That's downtime by design. It exists for apps where two versions must never run at the same time. - New pods count as available too early. A rolling update removes an old pod once a new one is available. Without a readiness probe, a pod is Ready the moment its container starts, even if the app inside needs 40 seconds to warm up. The Service sends it traffic at once, and the rollout moves on.
- Old pods die while traffic still reaches them. When a pod is deleted, the kubelet starts the shutdown while, at the same time, the control plane removes the pod from the EndpointSlices, and every node's kube-proxy and any ingress controller catch up a moment later. An app that exits as soon as it gets SIGTERM refuses the requests still in flight (the full deletion sequence).
Diagnosis and fix
1. Check the strategy
$ kubectl describe deploy storefront -n front | grep -E "StrategyType|RollingUpdate"
StrategyType: RecreateThere's no RollingUpdateStrategy line at all. Switch to a rolling update that never drops below the desired count:
$ kubectl patch deploy storefront -n front -p '{"spec":{"strategy":{"type":"RollingUpdate","rollingUpdate":{"maxSurge":1,"maxUnavailable":0}}}}'
deployment.apps/storefront patchedmaxSurge: 1 allows one extra pod during the rollout. maxUnavailable: 0 means an old pod is only removed after a new one is available. The defaults are 25% / 25%, and the rounding matters:
$ kubectl explain deployment.spec.strategy.rollingUpdate.maxUnavailable
...
The maximum number of pods that can be unavailable during the update. Value
can be an absolute number (ex: 5) or a percentage of desired pods (ex: 10%).
Absolute number is calculated from percentage by rounding down. This can not
be 0 if MaxSurge is 0. Defaults to 25%.With 3 replicas, 25% unavailable rounds down to 0 and 25% surge rounds up to 1. So the default is already safe for a 3-replica Deployment. Somebody had to choose Recreate.
2. Make "available" mean "can serve"
Give the container a readiness probe that checks the app's own readiness endpoint, not just "the process exists":
readinessProbe:
httpGet:
path: /ready
port: http
periodSeconds: 5Now a new pod joins the Service only once it answers, and the rollout waits for it. Make sure the path is right: a readiness probe that never passes keeps every pod out of the Service. For slow-starting apps like a JVM, add a startupProbe rather than a long liveness delay. A liveness probe that fires before the app is up turns every start into a restart, and eventually into CrashLoopBackOff.
3. Close the endpoint-removal race
Delay the SIGTERM until routing has caught up, then let the app drain:
spec:
terminationGracePeriodSeconds: 30
containers:
- name: web
lifecycle:
preStop:
sleep:
seconds: 5The built-in sleep action needs Kubernetes 1.30 or newer. On older clusters use exec: {command: ["sleep", "5"]}, which needs a sleep binary in the image. During those 5 seconds the pod keeps serving while it's removed from every node's rules. Then SIGTERM arrives, and the app should stop accepting new connections and finish the ones in flight (in Spring Boot that's server.shutdown=graceful, the default since 3.4). The arithmetic has to fit:
preStop sleep 5 + in-flight drain (say 20 s) + a few seconds < terminationGracePeriodSeconds 30The preStop time counts towards the grace period. If the sum is too big, the kubelet sends SIGKILL mid-drain, and you're back to dropped requests, plus exit code 137 in the pod status. If the app never drains at all, check that SIGTERM reaches it. A shell script as the container's PID 1 often swallows it (why containers need a PID 1 that forwards signals).
4. Prove it with one rollout
$ kubectl set image deploy/storefront -n front web=nginx:1.28
deployment.apps/storefront image updated
$ kubectl rollout status deploy/storefront -n front --timeout=120s
Waiting for deployment spec update to be observed...
Waiting for deployment "storefront" rollout to finish: 0 out of 3 new replicas have been updated...
Waiting for deployment "storefront" rollout to finish: 2 out of 3 new replicas have been updated...
Waiting for deployment "storefront" rollout to finish: 1 old replicas are pending termination...
deployment "storefront" successfully rolled outKeep a loop running against the Service while it rolls (while true; do curl -s -o /dev/null -w "%{http_code}\n" http://storefront/; sleep 0.2; done). Every line should be a 200.
How to prevent it
- Don't use
Recreateunless two versions really can't coexist, and write down why. - Give every serving container a readiness probe, and a startup probe if it's slow.
- Add a short
preStopsleep and graceful shutdown together. Most teams have one and not the other. - PodDisruptionBudgets don't help here. They limit evictions (drains), not Deployment rollouts.
maxUnavailableis what protects a rollout. - Watch the error rate during every deploy in a dashboard, not only in the uptime monitor.
Practise it
Chapter 15's incident "every deploy is a 30-second outage" starts from Recreate and asks you to prove the rollout never drops below 3 available pods. The probe lessons in chapter 17 and the shutdown lessons in chapters 3 and 21 cover the rest.