OnCallReady

KubernetesProcesses · 1 min read

What happens between kubectl delete pod and the container disappearing

SIGTERM, the grace period, SIGKILL - and the endpoint removal running at the same time. Why pods drop requests on deploy, and how preStop and graceful shutdown fix it.

kubectl delete pod web-7d9f-x2k4p returns at once, but a lot happens before the container is really gone, and two of those things happen at the same time.

The sequence

  1. The API server marks the pod for deletion (deletionTimestamp) with a grace period: terminationGracePeriodSeconds, 30 s by default. kubectl get pods shows it as Terminating.
  2. In parallel: · the kubelet on the node runs the preStop hook, if there is one, then sends SIGTERM to the container's main process (PID 1 in the container); · the endpoints controller removes the pod from the Service's endpoints, and every node's kube-proxy (and any ingress controller) updates its rules.
  3. When the grace period runs out and the process is still alive, the kubelet sends SIGKILL. The container exits with 137 (128 + 9).

Why deploys drop requests

Step 2 is a race. If the app exits the instant it gets SIGTERM, some kube-proxies and load balancers haven't caught up yet and still send it traffic, and those requests fail.

The usual fixes:

  • a short preStop sleep (a few seconds), so routing catches up before SIGTERM arrives;
  • graceful shutdown in the app: on SIGTERM stop accepting new connections, finish the requests in flight, then exit (Spring Boot: server.shutdown=graceful);
  • a grace period longer than preStop + the longest request.

The same race runs on every rolling update, once per old pod: rolling updates without downtime puts these fixes together with readiness probes.

PID 1 and signals

SIGTERM goes to PID 1 in the container. If that's a shell script that doesn't pass signals on (sh -c "java ..."), the app never hears it and gets SIGKILLed 30 s later. Use exec in entrypoint scripts, or a tiny init such as tini (why PID 1 has extra jobs).

It's the same pattern as systemd

systemctl stop does the same dance: SIGTERM, wait TimeoutStopSec (90 s by default), then SIGKILL. Learn it once on a Linux box and the Kubernetes version is just the same thing at cluster scale.

OnCallReady is free, with no ads and no tracking. RSS · All posts