Why this lesson
Every incident asks you to decide before you know. The page says checkout is slow and failing; a release went out 25 minutes ago; the database on-call is sure it is a missing index; the network team says nothing changed. You will not have the root cause for an hour. You have to choose something now, with people watching, knowing you may be wrong.
This lesson is how to decide well under that pressure: what to do first, which kinds of change to prefer, how to keep a decision honest with a time box, and how to write it down so the next person (and the postmortem) knows why.
What you need to know already: 0.29 (mitigate first, the mitigation menu), 37.4 (polling for strong objections), 37.14 (escalation).
Mitigate first, understand later
Google's incident guidance puts the priorities in one line: "Stop the bleeding, restore service, and preserve the evidence for root-causing." The order matters. Users are hurt now; the root cause can be found tomorrow, with time and without pressure. The SRE Workbook calls the first moves generic mitigations - actions that help for many causes without knowing which one it is: roll back the last change, drain traffic away from a bad region or copy, turn off a feature flag, add capacity.
The question is not "what is broken?" but "what is the cheapest action that makes users hurt less, that we can undo if we are wrong?" Chapter 0's menu, from cheapest to most drastic, still holds: roll back, flag off, drain or fail over, scale out, shed load, restart (and capture evidence before a restart destroys it - a thread dump, the logs, the heap).
The rollback bias
If a change went out shortly before the impact started, roll it back first. You do not need to prove the change caused it. Most incidents follow a change, and a rollback is usually the fastest action with a known result. That default has a name: the rollback bias.
It has limits, and a good IC asks about them before saying yes:
- Is the rollback safe? A release that migrated the database schema forward may not run backwards: if the new code wrote data in a new shape, the old code may choke on it. Ask the person who owns the change: "if we roll back, what breaks?"
- Has it been done before? A rollback path nobody has used in a year is itself a change. In chapter 0's #4471, the runbook's rollback command no longer existed and cost 13 minutes.
- Is the timing really a match? A release at 15:05 and errors from 15:06 is a strong hint. A release last Tuesday is not.
One-way and two-way doors
Amazon's 2015 shareholder letter split decisions into two types: one-way doors (consequential and hard or impossible to reverse - walk through and you cannot come back) and two-way doors (if it is wrong, you walk back through). The letter argues that two-way doors should be decided fast by small groups, and only one-way doors deserve slow, careful deliberation.
In an incident, that is your filter:
two-way door (decide fast) one-way door (stop and think, escalate)
----------------------------------- -----------------------------------------
roll back a code release restore a database from backup (loses recent data)
turn a feature flag off delete or rewrite data
drain traffic from one region a schema change in production under load
scale out failing over in a way that cannot fail back
pause a marketing send make a public statement about the cause
The trap is a one-way door that looks like a quick fix. A CREATE INDEX on a big table "takes a few minutes" - and, on PostgreSQL without CONCURRENTLY, blocks every write to that table while it runs. Under pressure people reach for exactly that kind of change, because it feels like progress. Say no, or not yet, and say why.
OODA: a loop, not a plan
US Air Force Colonel John Boyd described how pilots (and later, organisations) win under uncertainty as a loop: Observe, Orient, Decide, Act (OODA) - and whoever runs the loop faster and more accurately than the situation changes stays ahead of it. An incident maps onto it closely:
observe the impact now (the card, the dashboard), what people report
orient what changed, what we know, the hypotheses still alive, what would prove
each one wrong; this is where experience lives
decide one action, with an owner, a time box and an abort condition
act the ops lead does it; the IC says it out loud in the channel
...then observe again: did it do what we expected?
Two habits keep the loop honest. List hypotheses, not one theory: fixation (being sure early and reading every graph as proof) is the most common way smart responders lose an hour. And test the cheapest one first, not the most interesting one. In Richard Cook's words, from How Complex Systems Fail: "All practitioner actions are gambles." Make small gambles you can take back.
A decision that survives the incident
Write every significant decision down, with its reason, its owner, and how you will know it worked:
DECISION 15:34 roll back checkout to 3.0.1 (owner @dana)
because errors and latency started with 3.0.2 at 15:05; Mo confirms the only
migration adds a nullable column, so 3.0.1 runs fine on it
check p99 under 800 ms and errors under 1% within 10 min of the rollback
if not we look at the database next (Priya's index theory), with a plan that
does not lock the table
The time box and the abort condition ("if not") are what make a decision safe to take with half the facts: you have agreed in advance when you will stop waiting for it to work. In the room, incident decide records the decision and its why on the timeline. This lesson's room is INC-4529 - the slow checkout from the introduction:
$ incident log
15:30 bot PagerDuty: [FIRING] CheckoutLatencyBurn severity=page
15:30 bot checkout: p99 4.1 s (SLO 800 ms), 6.2% of requests failing. Release 3.0.2 went out at 15:05.
$ incident say "@mo what changed? is a rollback of 3.0.2 safe?"
15:31 @mo 3.0.2 went out at 15:05 (order history on the checkout page, one new query per checkout). It has one migration, but it only adds a nullable column - 3.0.1 ignores it. Rolling back the code is safe; the column can stay.
15:31 @priya Probably the database: a query on orders by customer_id is doing a full scan. I can add the index right now in production, a few minutes.
$ incident say "@priya hold the index for now - a CREATE INDEX on that table locks writes"
15:32 @priya Fair - on a table this size CREATE INDEX without CONCURRENTLY takes a lock. I will prepare it properly for tomorrow.
15:32 @lee Network team says nothing changed on their side today, for what it is worth.
$ incident decide "roll back checkout to 3.0.1" --why "errors and latency started with 3.0.2 and the rollback is safe; if p99 is not under 800 ms 10 min after it finishes, we look at the database next"
Decision recorded at 15:32 UTC.
$ incident say "Proposal: roll back to 3.0.1. Any strong objections?"
15:34 @dana No objection from me.
Asked, checked the door, recorded, polled. Total cost: four minutes. PagerDuty's IC guide again: "Making the 'wrong' decision is better than making no decision." A wrong two-way-door decision costs you ten minutes and teaches you something. No decision costs you the whole hour.
When you cannot decide
Sometimes the decision is not yours to make: taking the whole shop offline, spending on emergency capacity, telling customers something legally sensitive. That is hierarchical escalation (37.14): page the person who can decide, and give them a proposal, not a problem - "I recommend we take checkout offline for 20 minutes to stop duplicate charges; I need your yes in 5 minutes".
What you can now do
- put mitigation before diagnosis, and choose generic mitigations first
- apply the rollback bias, and check its limits (is it safe, has it been done, does the timing match)
- sort actions into two-way and one-way doors, and refuse risky one-way changes under pressure
- run the OODA loop with several hypotheses and the cheapest test first
- record a decision with its reason, owner, time box and abort condition