OnCallReady

Lesson 3.18 · Processes & Signals · 8 min read

A graceful stop: SIGTERM, grace period, SIGKILL

In plain words

Imagine a shop closing for the evening. Two things should happen: the sign on the door turns to "closed", and the manager tells the cashier to finish up. If the manager speaks first, people who were already walking towards the shop still come in for a few seconds, because they have not seen the sign. If the cashier locks the till the instant they are told, those last customers are turned away angry. And if the cashier never hears the manager at all, security eventually drags everyone out mid-purchase.

That is a graceful stop. "Finish up" is SIGTERM. The time the cashier gets is the grace period (TimeoutStopSec in systemd). Security is SIGKILL. The sign is the load balancer that sends customers, and a few seconds of patience after SIGTERM covers the people already on their way.

Why every deploy causes a few errors

Every time a new version of a service is installed, the old one has to stop. If it stops badly, the requests it was handling fail - a small spike of errors on every deploy that nobody can reproduce afterwards. This lesson is the exact sequence of a graceful stop, and the two ways it goes wrong.

What you need to know already: 2.24 (systemctl stop: ExecStop, SIGTERM, TimeoutStopSec, SIGKILL), 3.6 (signals and handlers).

The sequence, on any supervisor

A supervisor is the program that starts and stops your service - on this box, systemd. Every supervisor stops a service the same way:

1. (optional) run a stop hook first            systemd: ExecStop=
2. send SIGTERM                                 "please finish up"
3. wait up to a GRACE PERIOD                    systemd: TimeoutStopSec= (90s)
4. SIGKILL whatever is still alive              "stop now"

The grace period is the time the program gets between SIGTERM and SIGKILL. A well-behaved program uses it to drain: stop accepting new work, finish the requests already in flight, close connections, then exit on its own. A program that exits early frees the service at once; the grace period is a maximum, not a delay.

Failure 1: the program ignores SIGTERM

No handler, a handler that does nothing, or a wrapper script that receives the SIGTERM and never passes it to the real program (a shell running a script does not forward signals to the command it is waiting for). Then:

The fix for the wrapper case: end the script with exec ./the-real-program. exec replaces the shell with the program, so the program keeps the same PID and gets the SIGTERM itself.

Failure 2: traffic keeps arriving after SIGTERM

Services with more than one copy usually sit behind a load balancer: a program that receives all incoming requests and spreads them over the copies. When one copy is stopped, two things have to happen:

If (b) starts before (a) has fully taken effect, a copy that already received SIGTERM still gets new requests for a second or two. If it responds to SIGTERM by closing its listening socket at once, those requests get "connection refused". That is the error spike during a deploy.

The standard fix: on SIGTERM, keep serving for a few seconds (or have the stop hook just sleep 5 before the SIGTERM is sent), so the load balancer has stopped routing to you before you stop listening. It looks like a hack. It is the documented pattern.

How it ended: the exit code tells you

When a process is ended by a signal, the shell and supervisors report the exit code 128 + the signal number:

143 = 128 + 15   SIGTERM: a normal stop, if it exited on it
137 = 128 + 9    SIGKILL: the grace period ran out - or the kernel killed it
                 for using too much memory (Chapter 5)
130 = 128 + 2    SIGINT: someone pressed Ctrl+C

So a service that always exits with 137 on a normal stop is telling you it ignores SIGTERM.

Both failures have one root cause: an application that does not handle SIGTERM properly. The only symptom is a small error spike correlated with deploys, which nobody attributes to the deploy.

Later (Ch 15): Kubernetes stops a pod (its unit of running software) with the same four steps: kubectl delete pod -> the preStop hook (= ExecStop) -> SIGTERM to the container's PID 1 -> wait terminationGracePeriodSeconds (default 30s, = TimeoutStopSec) -> SIGKILL, exit

  1. Removing the pod from the Service's endpoints (= taking it out of the load

balancer) happens at the same time as the SIGTERM, not before it - which is exactly failure 2, and why a preStop sleep of ~5 seconds is the standard fix.

What you can now do

Why it helps

Every deploy stops the old version of a service. An app that closes its listener the instant it gets SIGTERM, or never receives SIGTERM because a wrapper script swallowed it, drops requests on every release, and the only symptom is a small error spike nobody links to the deploy.

Knowing the sequence lets you fix it properly: SIGTERM handling that stops accepting and drains, a short delay so the load balancer stops routing first, a grace period longer than your slowest request, and exec in wrapper scripts. It also explains exit codes 143 and 137 on stopped services, and why a stop that always takes exactly 90 seconds means the app ignores SIGTERM. The same sequence comes back in every tool that runs services, so it is a favourite interview question.

Commands in this lesson

sleep

FAQ

What is the difference between the grace period and a delay?

The grace period is a maximum, not a wait. systemd sends SIGTERM and then waits up to TimeoutStopSec= (90 seconds by default). A program that drains and exits after 3 seconds frees the service after 3 seconds. Only a program that ignores SIGTERM, or takes too long, uses the whole period and then gets SIGKILL.

Why would requests still arrive after SIGTERM?

Because the thing sending requests, usually a load balancer in front of several copies of the service, finds out about the stop separately and a little later. For a second or two it keeps routing to a copy that is already shutting down. If that copy stops listening at once, those requests fail. Keeping the listener open for a few seconds after SIGTERM covers the gap.

Does the stop hook's time count against the grace period?

In systemd, ExecStop= runs first and has its own timeout, then SIGTERM is sent and the grace period applies to what is left of the unit. The practical rule is the same everywhere: size the grace period as any pre-stop delay plus the longest drain the app needs, with a margin, and measure how long a real stop takes under load.

What if the program ignores SIGTERM?

It keeps running until the grace period ends, then the supervisor sends SIGKILL and it exits with 137. Every stop takes the full timeout, deploys slow down, and in-flight requests are cut. Common causes are a shell script that started the program without exec, so the signal never reaches it, or a program with no SIGTERM handler that happens to ignore it.

How do I see which case I am in?

Time a stop: time sudo systemctl stop unit. A stop that takes exactly TimeoutStopSec hit the timeout. The journal says it too: State 'stop-sigterm' timed out. Killing. and Failed with result 'timeout'. systemctl show -p ExecMainStatus unit gives the exit code; 143 (or 0) is a clean TERM exit, 137 a SIGKILL.

In an interview Junior

Users see a few errors on every deploy. What is probably going on?

Something is wrong with the graceful stop. Every supervisor stops a service the same way: an optional stop hook (ExecStop=), SIGTERM, a grace period (TimeoutStopSec=, 90 s by default), then SIGKILL. A well-behaved program uses the grace period to drain: stop taking new work, finish the requests in flight, exit.

Two usual failures:

  1. It ignores SIGTERM - no handler, or a wrapper script that never passes the signal on. Every stop takes the full grace period and ends in SIGKILL, cutting requests off. The tell: it always exits 137 instead of 143. Fix: handle SIGTERM, and end wrapper scripts with exec ./the-real-program.
  2. Traffic still arrives after SIGTERM - the load balancer has not stopped routing to it yet, and it closes its socket at once: "connection refused". Fix: keep serving a few seconds after SIGTERM (or sleep 5 in the stop hook).

Also asked: What do exit codes 143, 137 and 130 tell you about how a process ended? · What is a grace period, and how would you choose its length? · Why does a shell script wrapper break a graceful stop, and what does exec fix?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.