Imagine a toy car that keeps crashing into the wall. A helpful parent picks it up and restarts it each time. But after five crashes in one minute, a sensible parent says "this car is broken" and puts it on the shelf until someone fixes it, instead of restarting it all night.
That rule is the start limit: StartLimitBurst=5 starts within StartLimitIntervalSec=60, then failed (Result: start-limit-hit). The trap is where you write the rule: it belongs in [Unit]; in [Service] it is ignored with only a warning. The second trap: if the parent waits ten seconds between restarts (RestartSec=10) but only counts crashes in a ten-second window, the rule can never trigger. reset-failed takes the car off the shelf.
Why this matters
A service with Restart=always that crashes the instant it starts will be restarted, crash, be restarted... hundreds of times a minute, filling the log and burning CPU, while looking "up" half the time. This endless cycle is a crash loop. systemd's start limit is the brake: after too many starts in a short window, it gives up and marks the unit failed so a human looks at it.
What you need to know already: 2.10 (Restart=, SIGKILL, kill -9), 2.7 (systemd-analyze verify).
Rate limiting stops a crash loop eating the box
[Unit]
StartLimitIntervalSec=60
StartLimitBurst=5
"More than 5 starts within 60 seconds and I give up." The unit goes to:
Active: failed (Result: start-limit-hit)
and stays there until a human intervenes. Without this, a service that crashes instantly with Restart=always would spin forever.
StartLimitIntervalSec=0 disables the limit entirely - restart forever, never give up. Some agent programs ship like that on purpose: the reason they fail is often outside themselves (a setting on the machine, another server not up yet), so they should keep trying until it is fixed. You will meet one in 2.36.
The gotcha: they go in [Unit], not [Service]
This is the one that cost you time on day one. StartLimitIntervalSec and StartLimitBurst are [Unit] directives. Put them in [Service] and systemd ignores them - with a warning you only see if you go looking:
$systemd-analyze verify /etc/systemd/system/legacy-demo.service
.../legacy-demo.service:9: Unknown key name 'StartLimitIntervalSec' in section 'Service', ignoring.
systemctl daemon-reload prints nothing. systemctl status looks fine. The unit just silently has no rate limit.
The second gotcha: the defaults fight each other
The default is StartLimitIntervalSec=10s, StartLimitBurst=5. If you also set RestartSec=10, each restart is 10 seconds apart - so the 10-second window never contains more than one start, and the limit can never trip. The service crash-loops forever while looking correctly configured.
Either widen the interval (60s with RestartSec=10 gives you a real limit of ~5) or shorten RestartSec. The two numbers only mean something together.
Getting out of failed
Once a unit hits the limit it stays failed and systemd will not start it again by itself.
systemctl list-units --failed # what is broken right now
systemctl reset-failed demo # clear the failed state AND the counter
sudo systemctl start demo
list-units --failed shows only units in the failed state, with the same columns as before (UNIT LOAD ACTIVE SUB DESCRIPTION). reset-failed with no argument clears every failed unit. It does not fix anything - it just lets you try again.
What you can now do
put the start limit in the right section and choose numbers that can trip
spot a failed unit box-wide and clear it after fixing the cause
Why it helps
Start limits decide what a broken deploy looks like: either the service gives up and fails visibly (so an alert can fire), or it restart-loops forever, filling the journal and hammering whatever it depends on. Both can be right. Some agents deliberately set StartLimitIntervalSec=0 so they never give up - which is why the incident at the end of this chapter shows a service stuck in activating (auto-restart) for hours.
This lesson is also a real review catch: StartLimit* in [Service] is common in old blog posts and old units, and systemd-analyze verify is the only thing that tells you. And knowing reset-failed is how you bring a service back right after a fix, without waiting.
Older systemd versions had StartLimitInterval= and StartLimitBurst= in [Service]. They moved to [Unit] (and StartLimitInterval was renamed StartLimitIntervalSec) because the limit applies to every unit type, not just services. That is why old tutorials still show the old place. With the modern names in [Service], systemd ignores them with a warning. Put them in [Unit].
Does the start limit count manual starts too?
Yes. Every start attempt within the window counts, whether it came from Restart=, from you, or from another unit. Running systemctl restart several times quickly while testing can trip the limit yourself, and you will see "start request repeated too quickly". systemctl reset-failed unit clears the counter so you can try again.
What happens after start-limit-hit? Will it ever retry?
Not by itself. The unit stays failed until someone runs systemctl reset-failed and starts it, or something else starts it after the window has passed. StartLimitAction= can choose a different reaction, like rebooting the machine, which small appliances sometimes do. On servers, pair the limit with OnFailure= (2.28) so a human gets told.
Is StartLimitIntervalSec=0 a bad idea?
It switches the limit off, so the service retries forever. That is right for agents whose failures usually come from outside and fix themselves - waiting for a config file, a certificate or the network. It is wrong where each failed attempt does damage or makes noise, like a job that corrupts data each time it starts. If you switch it off, choose a long enough RestartSec so it does not spin.
How do I find every failed unit on a box?
systemctl list-units --failed, or the short form systemctl --failed, lists units in the failed state. systemctl is-system-running sums it up: degraded means at least one unit has failed. After fixing, systemctl reset-failed with no unit name clears all of them. Monitoring tools often simply alert on the degraded state.
In an interview Junior
A service shows "start request repeated too quickly". What does it mean, and what do you do?
It hit its start limit: more than StartLimitBurst starts within StartLimitIntervalSec (default 5 in 10 seconds). systemd gave up and marked it failed (Result: start-limit-hit); it will not start again by itself - that is the brake on a crash loop.
What I do: systemctl status svc for the last exit code, then journalctl -u svc for the program's own error, usually the same line over and over (a missing config file, a port already in use). Fix that, then sudo systemctl reset-failed svc and start it again. reset-failed fixes nothing by itself; it only lets you try.
Two gotchas: the StartLimit* settings belong in [Unit] - in [Service] they are ignored with only a warning from systemd-analyze verify. And with RestartSec=10 the default 10-second window can never hold five starts, so the limit never trips.
Also asked: What is a crash loop, and how does systemd stop one? · How do you list every failed unit on a box? · What does StartLimitIntervalSec=0 do, and when would a service want it?