What happens on SIGTERM
The problem. Every deploy and every node drain stops pods. If the app dies in the middle of a request, users see errors "only during deploys" - which on a busy day means all the time. Graceful shutdown finishes the work in hand first, as long as the outer timeout lets it.
What you need to know already: SIGTERM, SIGKILL and exit 143 (3.6), systemd TimeoutStopSec (2.24), what happens when a pod is deleted (3.18), readiness (17.20), EndpointSlices and kube-proxy (16.1, 16.3).
The JVM receives SIGTERM, runs its shutdown hooks (code the app registered to run on exit), and Spring closes the application context (the container of all the app's beans, 21.1). With graceful shutdown the web server first stops accepting new requests, waits for in-flight ones to finish, and only then do the other beans (the DataSource, schedulers) shut down.
server.shutdown=graceful the DEFAULT since Spring Boot 3.4 (before: immediate)
spring.lifecycle.timeout-per-shutdown-phase=30s how long to wait for in-flight requests (default 30s)
In the log:
o.s.b.w.e.tomcat.GracefulShutdown : Commencing graceful shutdown. Waiting for active requests to complete
o.s.b.w.e.tomcat.GracefulShutdown : Graceful shutdown complete
com.zaxxer.hikari.HikariDataSource : HikariPool-1 - Shutdown initiated...
com.zaxxer.hikari.HikariDataSource : HikariPool-1 - Shutdown completed.
And when in-flight work outlives the phase timeout:
o.s.c.support.DefaultLifecycleProcessor : Shutdown phase 2147482623 ends with 1 bean still running after timeout of 30000ms: [webServerGracefulShutdown]
o.s.b.w.e.tomcat.GracefulShutdown : Graceful shutdown aborted with one or more requests still active
Readiness flips to REFUSING_TRAFFIC (OUT_OF_SERVICE, 503) as soon as shutdown starts, so a load balancer that probes readiness stops routing to it.
The JVM then exits with 143 (128 + 15, SIGTERM). That is why our units say SuccessExitStatus=143 - otherwise systemd would call a normal stop a failure.
The outer timeout always wins
Whoever sent SIGTERM has its own deadline, after which it sends SIGKILL:
systemd TimeoutStopSec= default 90s (DefaultTimeoutStopSec in system.conf)
Kubernetes terminationGracePeriodSeconds default 30s
Docker docker stop -t default 10s
If the application's graceful phase is longer than the outer timeout, the process is killed mid-drain: in-flight requests get a closed connection (curl: (52) Empty reply from server, a 502 at the proxy), and in systemd:
orders.service: State 'stop-sigterm' timed out. Killing.
orders.service: Killing process 1210 (java) with signal SIGKILL.
orders.service: Main process exited, code=killed, status=9/KILL
orders.service: Failed with result 'timeout'.
The rule: timeout-per-shutdown-phase + everything else in shutdown < the outer timeout. In Kubernetes with a preStop sleep (a hook the kubelet runs before sending SIGTERM):
preStop sleep 5 + phase timeout 20s + a few seconds for the rest < terminationGracePeriodSeconds 30
The preStop trap
When a pod is deleted, two things happen concurrently, not in order:
kubelet control plane
runs preStop, then sends SIGTERM removes the pod from Endpoints / EndpointSlices
kube-proxy on every node updates iptables/IPVS
ingress controllers update their upstreams
The endpoint removal takes a moment to reach every node. If the app stops accepting connections the instant SIGTERM arrives, traffic that is still being routed to it fails - 502s on every deploy (502 Bad Gateway = the proxy in front could not get an answer, 9.23). The fix is to delay the SIGTERM:
lifecycle:
preStop:
exec: { command: ["sleep", "5"] } # or the built-in: preStop: { sleep: { seconds: 5 } } (1.30+)
The pod keeps serving for 5 seconds while the endpoint removal propagates, then graceful shutdown drains whatever is still in flight. Two fixes for "502s on every deploy": the preStop delay, and graceful shutdown with a phase timeout that fits the grace period. Most teams have one and not the other.
On this box
systemd plays the kubelet's role: systemctl stop sends SIGTERM and waits TimeoutStopSec. The same arithmetic applies, and you can watch it:
$ curl -s -o /dev/null -w '%{http_code}\n' localhost:8080/api/reports/export & # a 45-second request
$ sudo systemctl restart orders
$ journalctl -u orders -n 8 --no-pager
Next mission: a phase timeout of 120 s against a TimeoutStopSec of 30 s.
What you can now do
- Turn on graceful shutdown and read its log lines.
- Check the arithmetic: preStop + phase timeout + the rest < TimeoutStopSec or terminationGracePeriodSeconds.
- Explain why a preStop sleep stops the 502s on every deploy.