OnCallReady

systemdLinux · 1 min read

systemd: your StartLimitIntervalSec is in the wrong section (and silently ignored)

StartLimitIntervalSec and StartLimitBurst belong in [Unit]. In [Service] systemd ignores them with a warning you only see in systemd-analyze verify. Plus the default that never trips.

You wrote a service with Restart=always, and to stop a crash loop from eating the box you added a start limit:

ini
[Service]
ExecStart=/usr/local/bin/demo.sh
Restart=always
RestartSec=10
StartLimitIntervalSec=60
StartLimitBurst=5

systemctl daemon-reload prints nothing and systemctl status looks fine. The service then crash-loops forever.

The limit lives in [Unit]

StartLimitIntervalSec= and StartLimitBurst= are [Unit] settings. In [Service], StartLimitIntervalSec is an unknown key, and systemd ignores it with a warning that you only see if you go looking:

terminal
$ systemd-analyze verify /etc/systemd/system/demo.service
.../demo.service:6: Unknown key name 'StartLimitIntervalSec' in section 'Service', ignoring.

The fix:

ini
[Unit]
StartLimitIntervalSec=60
StartLimitBurst=5

[Service]
ExecStart=/usr/local/bin/demo.sh
Restart=always
RestartSec=10

Then sudo systemctl daemon-reload. Make systemd-analyze verify part of writing a unit: it also catches typos and ExecStart paths that don't exist.

The second gotcha: defaults that cancel out

Without your own values, the default is 5 starts in 10 seconds:

terminal
$ systemctl show -p StartLimitBurst -p StartLimitIntervalUSec ssh
StartLimitBurst=5
StartLimitIntervalUSec=10s

With RestartSec=10, each restart comes 10 seconds after the last, so a 10-second window never holds more than one start. The limit can never trip. Either widen the interval (say 60 s for 5 starts, as above) or shorten RestartSec.

When it does trip

output
Active: failed (Result: start-limit-hit)

The unit stays failed until a human looks: systemctl list-units --failed shows it, and sudo systemctl reset-failed demo clears the counter after you've fixed the cause. Kubernetes has its own version of this back-off: CrashLoopBackOff.

StartLimitIntervalSec=0 disables the limit completely. That's right for agents whose failures are usually caused by something outside them (like kubelet waiting for the machine to be set up properly). It's wrong for most application services. Docker's --restart always has the same blind spot: a container that keeps restarting.

OnCallReady is free, with no ads and no tracking. RSS · All posts