Why this lesson
The on-call week ends and the team's planned work has not moved: every day went on restarting the same program, cleaning the same disk, renewing the same kind of certificate by hand. "On-call is eating the sprint" (the sprint is the team's two-week block of planned work), someone says, and nothing changes, because it is a feeling. This lesson turns it into a measured, ranked list of what to remove.
What you need to know already: 0.30 Blameless postmortems, toil (the six properties, overhead vs engineering).
Measure before you argue
"On-call is eating the sprint" is an opinion. "64% of on-call time last week was toil, and one category was a third of it" is a plan. Toil lives in the ticket system, the pager and the chat; you measure it the same way as everything else in this chapter - count, group, sum.
$ cd ~/oncall-lab/labs/0-sre
$ awk -F'\t' '!/^#/ {m[$3]+=$4; n[$3]++} END {for (k in m) print m[k], n[k], k}' toil/oncall-week.tsv | sort -rn
288 14 restart-worker
209 6 disk-cleanup
181 1 postmortem
147 5 handover
121 2 runbook-update
119 3 cert-renewal
108 9 access-request
79 5 dns-change
The file has one ticket per line; column 3 is its category and column 4 the minutes spent. awk adds up the minutes (m) and counts the tickets (n) per category; sort -rn puts the biggest first. So each output line is: minutes, number of tickets, category.
Read it with the toil test in mind: manual, repetitive, automatable, interrupt-driven, no lasting value, grows with the service. restart-worker passes every test. postmortem is one ticket, not repetitive, and leaves the system better: engineering. handover is a meeting: overhead.
The toil budget
Google caps toil at 50% of an SRE's time. The number is less important than the reason: toil grows linearly with the service (twice the customers, twice the DNS tickets - DNS is the internet's address book that turns names like shop.example into server addresses), engineering does not. A team above the cap stops improving anything, then needs a new hire for every growth step. Track it per week, per person, and treat a rising trend as an incident in slow motion.
Toil is also a morale and risk problem, not just a time problem: repetitive manual changes are where typing mistakes happen, and people burn out on work that never ends.
Is it worth automating?
Automation costs time too. A quick way to rank candidates is hours per year:
$ awk 'BEGIN {printf "%-26s %8s %8s %10s\n", "task", "min/run", "runs/wk", "h/year"; n=split("restart-worker:20:14 dns-change:16:5 access-request:12:9 cert-renewal:38:1", t, " "); for (i=1; i<=n; i++) {split(t[i], f, ":"); printf "%-26s %8d %8d %10.0f\n", f[1], f[2], f[3], f[2]*f[3]*52/60}}'
task min/run runs/wk h/year
restart-worker 20 14 243
dns-change 16 5 69
access-request 12 9 94
cert-renewal 38 1 33
(split(s, arr, sep) cuts a text into a list and returns how many pieces - handy for small inline tables. Minutes per run x runs per week x 52 weeks / 60 = hours per year.) 243 hours a year is six working weeks of one person restarting one worker. Anything that costs less than that to fix pays back inside the year.
The table undercounts in three ways, all of which argue for automating sooner: the interrupt cost (a 20-minute task at 3am costs the whole night's sleep and the next morning), the growth (these numbers double when the service does), and the risk of a manual step going wrong. It overcounts nothing.
cert-renewal looks cheap at 33 hours - until the one renewal that is missed and the site's certificate expires. #4488 (a later mission) was a missed renewal. Frequency is not the only axis: a rare manual task with a catastrophic failure mode is a priority too.
Eliminate, then automate, then speed up
The order matters:
- Eliminate - remove the reason the task exists. payments-worker crashes because it loses its connection to the queue it reads work from: fix that and there is nothing to restart. A disk that fills because of overly chatty logging: turn the logging down.
- Automate - what must stay, a machine does: the system restarts the program itself when it crashes, a scheduled job runs the cleanup every night, certificates renew themselves, people request access through a form that grants it with an expiry date, DNS records are generated from a list.
- Speed up - only for what can be neither: a runbook with copy-pasteable commands, a script (a file of commands run in one go) that does the ten steps in one.
Automation that skips step 1 is a trap. Restarting a worker automatically every time it crashes hides a crash loop (a program that crashes, restarts, crashes again, forever): the toil is gone, the bug is not, and a half-processed batch on every crash might be a data problem nobody is looking at. Automate the symptom as a stopgap, and file the elimination as engineering work with an owner - the toil-week incident does exactly this.
Automation that is safe
Automation runs without a human watching, so it must fail safe:
idempotent running it twice does no harm (find -mtime +3 -delete is; "delete the
3 oldest files" is not)
bounded it cannot delete everything: -maxdepth 1, a name pattern, a limit
dry run it can print what it would do (find ... -print before -delete)
logged it says what it did (journalctl -u clean-exports)
observable it has its own alert if it stops running or starts failing
(a scheduled job that silently stops is worse than the toil it replaced)
stoppable one command disables it (systemctl disable --now clean-exports.timer)
Idempotent means "the same result however many times you run it". The examples use find, which searches for files: find /var/tmp -maxdepth 1 -type f -name 'checkout-export-*.csv' -mtime +3 -print -delete looks only in /var/tmp itself (-maxdepth 1: not in folders inside it), only at files (-type f) whose name matches the pattern (* = anything), last changed more than 3 whole days ago (-mtime +3), prints each one and deletes it. It is bounded, idempotent and records every file it removes. The toil mission next turns it into a scheduled job. (systemctl is the command that starts, stops and schedules programs on this kind of server; journalctl -u NAME shows the journal messages of one of them. The mission explains both.)
What is not toil
- Overhead: meetings, planning, 1:1s, training, HR, interviews. Necessary, not operational, and not something automation removes.
- Engineering: design, automation, SLOs, postmortems, fixing root causes, a painful one-off migration. Unpleasant is not the same as toil: if the service is permanently better afterwards, it was engineering.
- Incident response is operational work that needs judgement. The repetitive, scriptable parts of it (restarting the same thing by hand every time the same alert fires) are toil; deciding what is wrong is not.
In short
measure tickets and pages: count, group, sum minutes; track weekly
cap 50%: toil grows with the service, engineering does not
rank hours per year, plus interrupt cost, growth and risk
order eliminate > automate > speed up
safe idempotent, bounded, dry run, logged, observable, stoppable
What you can now do:
- measure toil from a ticket export and rank it by hours per year
- choose between eliminating, automating and speeding up a task
- judge whether a piece of automation is safe to leave running