Why every deploy causes a few errors
Every time a new version of a service is installed, the old one has to stop. If it stops badly, the requests it was handling fail - a small spike of errors on every deploy that nobody can reproduce afterwards. This lesson is the exact sequence of a graceful stop, and the two ways it goes wrong.
What you need to know already: 2.24 (systemctl stop: ExecStop, SIGTERM, TimeoutStopSec, SIGKILL), 3.6 (signals and handlers).
The sequence, on any supervisor
A supervisor is the program that starts and stops your service - on this box, systemd. Every supervisor stops a service the same way:
1. (optional) run a stop hook first systemd: ExecStop=
2. send SIGTERM "please finish up"
3. wait up to a GRACE PERIOD systemd: TimeoutStopSec= (90s)
4. SIGKILL whatever is still alive "stop now"
The grace period is the time the program gets between SIGTERM and SIGKILL. A well-behaved program uses it to drain: stop accepting new work, finish the requests already in flight, close connections, then exit on its own. A program that exits early frees the service at once; the grace period is a maximum, not a delay.
Failure 1: the program ignores SIGTERM
No handler, a handler that does nothing, or a wrapper script that receives the SIGTERM and never passes it to the real program (a shell running a script does not forward signals to the command it is waiting for). Then:
- every stop takes the full grace period,
- ends in SIGKILL, so no draining happened,
- and the in-flight requests are simply cut off.
The fix for the wrapper case: end the script with exec ./the-real-program. exec replaces the shell with the program, so the program keeps the same PID and gets the SIGTERM itself.
Failure 2: traffic keeps arriving after SIGTERM
Services with more than one copy usually sit behind a load balancer: a program that receives all incoming requests and spreads them over the copies. When one copy is stopped, two things have to happen:
- a) the load balancer must stop sending it new requests, and
- b) the copy must shut down.
If (b) starts before (a) has fully taken effect, a copy that already received SIGTERM still gets new requests for a second or two. If it responds to SIGTERM by closing its listening socket at once, those requests get "connection refused". That is the error spike during a deploy.
The standard fix: on SIGTERM, keep serving for a few seconds (or have the stop hook just sleep 5 before the SIGTERM is sent), so the load balancer has stopped routing to you before you stop listening. It looks like a hack. It is the documented pattern.
How it ended: the exit code tells you
When a process is ended by a signal, the shell and supervisors report the exit code 128 + the signal number:
143 = 128 + 15 SIGTERM: a normal stop, if it exited on it
137 = 128 + 9 SIGKILL: the grace period ran out - or the kernel killed it
for using too much memory (Chapter 5)
130 = 128 + 2 SIGINT: someone pressed Ctrl+C
So a service that always exits with 137 on a normal stop is telling you it ignores SIGTERM.
Both failures have one root cause: an application that does not handle SIGTERM properly. The only symptom is a small error spike correlated with deploys, which nobody attributes to the deploy.
Later (Ch 15): Kubernetes stops a pod (its unit of running software) with the same four steps:
kubectl delete pod-> thepreStophook (= ExecStop) -> SIGTERM to the container's PID 1 -> waitterminationGracePeriodSeconds(default 30s, = TimeoutStopSec) -> SIGKILL, exit
- Removing the pod from the Service's endpoints (= taking it out of the load
balancer) happens at the same time as the SIGTERM, not before it - which is exactly failure 2, and why a
preStopsleep of ~5 seconds is the standard fix.
What you can now do
- Describe a graceful stop step by step: stop hook, SIGTERM, grace period, SIGKILL.
- Explain why deploys drop requests, and the two fixes (
exec, a short delay). - Read exit codes 143, 137 and 130 as "which signal ended it".