OnCallReady

Lesson 0.35 · SRE Fundamentals · 12 min read

Measuring toil, and choosing what to remove first

In plain words

Imagine you have to empty the rubbish bin in your room every hour because it's tiny. You could keep emptying it forever. You could get a bigger bin that empties itself. Or you could ask why so much rubbish appears in the first place: maybe you're unwrapping sweets one by one, and buying them unwrapped would make the rubbish disappear entirely.

Toil reduction follows that order: eliminate the reason the task exists, then automate what must stay, then speed up what can be neither. You measure first: minutes and tickets per category from the ticket system, then hours per year (min/run x runs/week x 52 / 60), plus the interrupt cost, growth and risk. Automation must be safe: idempotent, bounded, dry-runnable, logged, observable and stoppable.

Why this lesson

The on-call week ends and the team's planned work has not moved: every day went on restarting the same program, cleaning the same disk, renewing the same kind of certificate by hand. "On-call is eating the sprint" (the sprint is the team's two-week block of planned work), someone says, and nothing changes, because it is a feeling. This lesson turns it into a measured, ranked list of what to remove.

What you need to know already: 0.30 Blameless postmortems, toil (the six properties, overhead vs engineering).

Measure before you argue

"On-call is eating the sprint" is an opinion. "64% of on-call time last week was toil, and one category was a third of it" is a plan. Toil lives in the ticket system, the pager and the chat; you measure it the same way as everything else in this chapter - count, group, sum.

$ cd ~/oncall-lab/labs/0-sre
$ awk -F'\t' '!/^#/ {m[$3]+=$4; n[$3]++} END {for (k in m) print m[k], n[k], k}' toil/oncall-week.tsv | sort -rn
288 14 restart-worker
209 6 disk-cleanup
181 1 postmortem
147 5 handover
121 2 runbook-update
119 3 cert-renewal
108 9 access-request
79 5 dns-change

The file has one ticket per line; column 3 is its category and column 4 the minutes spent. awk adds up the minutes (m) and counts the tickets (n) per category; sort -rn puts the biggest first. So each output line is: minutes, number of tickets, category.

Read it with the toil test in mind: manual, repetitive, automatable, interrupt-driven, no lasting value, grows with the service. restart-worker passes every test. postmortem is one ticket, not repetitive, and leaves the system better: engineering. handover is a meeting: overhead.

The toil budget

Google caps toil at 50% of an SRE's time. The number is less important than the reason: toil grows linearly with the service (twice the customers, twice the DNS tickets - DNS is the internet's address book that turns names like shop.example into server addresses), engineering does not. A team above the cap stops improving anything, then needs a new hire for every growth step. Track it per week, per person, and treat a rising trend as an incident in slow motion.

Toil is also a morale and risk problem, not just a time problem: repetitive manual changes are where typing mistakes happen, and people burn out on work that never ends.

Is it worth automating?

Automation costs time too. A quick way to rank candidates is hours per year:

$ awk 'BEGIN {printf "%-26s %8s %8s %10s\n", "task", "min/run", "runs/wk", "h/year"; n=split("restart-worker:20:14 dns-change:16:5 access-request:12:9 cert-renewal:38:1", t, " "); for (i=1; i<=n; i++) {split(t[i], f, ":"); printf "%-26s %8d %8d %10.0f\n", f[1], f[2], f[3], f[2]*f[3]*52/60}}'
task                        min/run  runs/wk     h/year
restart-worker                   20       14        243
dns-change                       16        5         69
access-request                   12        9         94
cert-renewal                     38        1         33

(split(s, arr, sep) cuts a text into a list and returns how many pieces - handy for small inline tables. Minutes per run x runs per week x 52 weeks / 60 = hours per year.) 243 hours a year is six working weeks of one person restarting one worker. Anything that costs less than that to fix pays back inside the year.

The table undercounts in three ways, all of which argue for automating sooner: the interrupt cost (a 20-minute task at 3am costs the whole night's sleep and the next morning), the growth (these numbers double when the service does), and the risk of a manual step going wrong. It overcounts nothing.

cert-renewal looks cheap at 33 hours - until the one renewal that is missed and the site's certificate expires. #4488 (a later mission) was a missed renewal. Frequency is not the only axis: a rare manual task with a catastrophic failure mode is a priority too.

Eliminate, then automate, then speed up

The order matters:

  1. Eliminate - remove the reason the task exists. payments-worker crashes because it loses its connection to the queue it reads work from: fix that and there is nothing to restart. A disk that fills because of overly chatty logging: turn the logging down.
  2. Automate - what must stay, a machine does: the system restarts the program itself when it crashes, a scheduled job runs the cleanup every night, certificates renew themselves, people request access through a form that grants it with an expiry date, DNS records are generated from a list.
  3. Speed up - only for what can be neither: a runbook with copy-pasteable commands, a script (a file of commands run in one go) that does the ten steps in one.

Automation that skips step 1 is a trap. Restarting a worker automatically every time it crashes hides a crash loop (a program that crashes, restarts, crashes again, forever): the toil is gone, the bug is not, and a half-processed batch on every crash might be a data problem nobody is looking at. Automate the symptom as a stopgap, and file the elimination as engineering work with an owner - the toil-week incident does exactly this.

Automation that is safe

Automation runs without a human watching, so it must fail safe:

idempotent    running it twice does no harm (find -mtime +3 -delete is; "delete the
              3 oldest files" is not)
bounded       it cannot delete everything: -maxdepth 1, a name pattern, a limit
dry run       it can print what it would do (find ... -print before -delete)
logged        it says what it did (journalctl -u clean-exports)
observable    it has its own alert if it stops running or starts failing
              (a scheduled job that silently stops is worse than the toil it replaced)
stoppable     one command disables it (systemctl disable --now clean-exports.timer)

Idempotent means "the same result however many times you run it". The examples use find, which searches for files: find /var/tmp -maxdepth 1 -type f -name 'checkout-export-*.csv' -mtime +3 -print -delete looks only in /var/tmp itself (-maxdepth 1: not in folders inside it), only at files (-type f) whose name matches the pattern (* = anything), last changed more than 3 whole days ago (-mtime +3), prints each one and deletes it. It is bounded, idempotent and records every file it removes. The toil mission next turns it into a scheduled job. (systemctl is the command that starts, stops and schedules programs on this kind of server; journalctl -u NAME shows the journal messages of one of them. The mission explains both.)

What is not toil

In short

measure     tickets and pages: count, group, sum minutes; track weekly
cap         50%: toil grows with the service, engineering does not
rank        hours per year, plus interrupt cost, growth and risk
order       eliminate > automate > speed up
safe        idempotent, bounded, dry run, logged, observable, stoppable

What you can now do:

Why it helps

"On-call is eating the sprint" is an opinion; "64% of on-call time was toil and restarting one worker was a third of it, 243 hours a year" is a plan that gets approved. You'll produce that analysis from ticket exports with awk. Situations: someone adds an automatic restart (Restart=on-failure) to a crashing worker and calls it fixed; you point out the crash loop and the half-processed batches are still there, and file the elimination work. A scheduled cleanup silently stopped months ago; you insist every automation has its own alert. cert-renewal looks cheap at 33 hours a year until the one missed renewal causes an outage. Interviewers ask "how do you decide what to automate?" to see this reasoning.

Commands in this lesson

cd awk

FAQ

How do I decide which toil to automate first?

Rank by hours per year (minutes per run times runs per week times 52, over 60), then adjust for what that undercounts: interrupt cost (a 20-minute task at 3am costs a night's sleep), growth (it doubles when the service does) and risk (manual steps are where mistakes happen). A rare task with a catastrophic failure mode, like certificate renewal, can be a priority despite low hours.

Why eliminate before automating?

Because automation can hide the real problem. An automatic restart on a worker that crashes every batch removes the restart toil but leaves the crash loop, and maybe half-processed batches nobody is watching. Fix the cause if you can (the missed broker heartbeat, the debug logging filling the disk). If you automate as a stopgap, file the elimination as engineering work with an owner.

What makes automation safe to run unattended?

Idempotent (running twice does no harm), bounded (it can't delete everything: one fixed folder, a name pattern, a limit), dry-runnable (it can print what it would do), logged (it records what it did), observable (it has its own alert if it stops running or fails), and stoppable (one command disables it). A scheduled job that silently stops is worse than the toil it replaced.

Is incident response toil?

The judgement part isn't: deciding what's wrong and what to do is engineering-level work. The repetitive, scriptable parts are, like manually restarting the same thing every time the same alert fires. That's a signal to automate the response and stop paging for it, or better, remove the cause.

How do I track toil over time?

From the ticket system and pager: count, group by category, sum minutes, weekly and per person, as the toil-week incident did with one awk command over an export. Classify each category against the toil test. Treat a rising trend as an incident in slow motion, since toil grows with the service while the team doesn't.

In an interview Junior

What is toil, and how is it different from other operational work?

Toil, as the SRE book defines it, is work that is manual, repetitive, automatable, tactical (interrupt-driven), has no enduring value - the service is no better afterwards - and grows with the service. Restarting the same crashing worker every night is the classic example.

It is not overhead (meetings, planning, training), which is necessary but not operational, and it is not engineering (automation, SLOs, postmortems, fixing a root cause, even a painful one-off migration). Unpleasant is not the same as toil: if the service is permanently better afterwards, it was engineering.

Why it matters: toil scales with the service and engineering does not, so Google caps SRE toil at 50% of the time. To choose what to remove first, I measure it from the ticket export, rank tasks by hours per year, and prefer eliminating a task over automating it.

Also asked: How do you decide whether a task is worth automating? · Why is "eliminate" better than "automate"? · What would you check before trusting a cleanup script that runs on a schedule?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.