OnCallReady

Lesson 31.21 · Incident Command & Communication · 15 min read

Deciding under uncertainty: mitigate first, rollback bias, OODA

In plain words

You smell burning in the kitchen. You do not stand there working out whether it was the oil, the heat or the timer; you take the pan off the heat. That is cheap, quick, and you can put it back if you were wrong. Throwing the whole dinner in the bin is different: once it is gone, it is gone, so that one deserves a moment's thought.

Incident decisions work the same way. Mitigate first, understand later: the cheapest action that makes users hurt less and that you can undo. If a change went out just before the trouble started, roll it back first (the rollback bias). Sort actions into two-way doors, which you decide fast, and one-way doors, which you stop and think about. Run a loop (observe, orient, decide, act), keep several hypotheses alive, and write each decision down with how you will know it worked and when you will give up on it.

Why this lesson

Every incident asks you to decide before you know. The page says checkout is slow and failing; a release went out 25 minutes ago; the database on-call is sure it is a missing index; the network team says nothing changed. You will not have the root cause for an hour. You have to choose something now, with people watching, knowing you may be wrong.

This lesson is how to decide well under that pressure: what to do first, which kinds of change to prefer, how to keep a decision honest with a time box, and how to write it down so the next person (and the postmortem) knows why.

What you need to know already: 0.29 (mitigate first, the mitigation menu), 37.4 (polling for strong objections), 37.14 (escalation).

Mitigate first, understand later

Google's incident guidance puts the priorities in one line: "Stop the bleeding, restore service, and preserve the evidence for root-causing." The order matters. Users are hurt now; the root cause can be found tomorrow, with time and without pressure. The SRE Workbook calls the first moves generic mitigations - actions that help for many causes without knowing which one it is: roll back the last change, drain traffic away from a bad region or copy, turn off a feature flag, add capacity.

The question is not "what is broken?" but "what is the cheapest action that makes users hurt less, that we can undo if we are wrong?" Chapter 0's menu, from cheapest to most drastic, still holds: roll back, flag off, drain or fail over, scale out, shed load, restart (and capture evidence before a restart destroys it - a thread dump, the logs, the heap).

The rollback bias

If a change went out shortly before the impact started, roll it back first. You do not need to prove the change caused it. Most incidents follow a change, and a rollback is usually the fastest action with a known result. That default has a name: the rollback bias.

It has limits, and a good IC asks about them before saying yes:

One-way and two-way doors

Amazon's 2015 shareholder letter split decisions into two types: one-way doors (consequential and hard or impossible to reverse - walk through and you cannot come back) and two-way doors (if it is wrong, you walk back through). The letter argues that two-way doors should be decided fast by small groups, and only one-way doors deserve slow, careful deliberation.

In an incident, that is your filter:

two-way door (decide fast)            one-way door (stop and think, escalate)
-----------------------------------   -----------------------------------------
roll back a code release              restore a database from backup (loses recent data)
turn a feature flag off               delete or rewrite data
drain traffic from one region         a schema change in production under load
scale out                             failing over in a way that cannot fail back
pause a marketing send                make a public statement about the cause

The trap is a one-way door that looks like a quick fix. A CREATE INDEX on a big table "takes a few minutes" - and, on PostgreSQL without CONCURRENTLY, blocks every write to that table while it runs. Under pressure people reach for exactly that kind of change, because it feels like progress. Say no, or not yet, and say why.

OODA: a loop, not a plan

US Air Force Colonel John Boyd described how pilots (and later, organisations) win under uncertainty as a loop: Observe, Orient, Decide, Act (OODA) - and whoever runs the loop faster and more accurately than the situation changes stays ahead of it. An incident maps onto it closely:

observe    the impact now (the card, the dashboard), what people report
orient     what changed, what we know, the hypotheses still alive, what would prove
           each one wrong; this is where experience lives
decide     one action, with an owner, a time box and an abort condition
act        the ops lead does it; the IC says it out loud in the channel
...then observe again: did it do what we expected?

Two habits keep the loop honest. List hypotheses, not one theory: fixation (being sure early and reading every graph as proof) is the most common way smart responders lose an hour. And test the cheapest one first, not the most interesting one. In Richard Cook's words, from How Complex Systems Fail: "All practitioner actions are gambles." Make small gambles you can take back.

A decision that survives the incident

Write every significant decision down, with its reason, its owner, and how you will know it worked:

DECISION 15:34  roll back checkout to 3.0.1 (owner @dana)
  because       errors and latency started with 3.0.2 at 15:05; Mo confirms the only
                migration adds a nullable column, so 3.0.1 runs fine on it
  check         p99 under 800 ms and errors under 1% within 10 min of the rollback
  if not        we look at the database next (Priya's index theory), with a plan that
                does not lock the table

The time box and the abort condition ("if not") are what make a decision safe to take with half the facts: you have agreed in advance when you will stop waiting for it to work. In the room, incident decide records the decision and its why on the timeline. This lesson's room is INC-4529 - the slow checkout from the introduction:

$ incident log
15:30  bot       PagerDuty: [FIRING] CheckoutLatencyBurn severity=page
15:30  bot       checkout: p99 4.1 s (SLO 800 ms), 6.2% of requests failing. Release 3.0.2 went out at 15:05.
$ incident say "@mo what changed? is a rollback of 3.0.2 safe?"
15:31  @mo       3.0.2 went out at 15:05 (order history on the checkout page, one new query per checkout). It has one migration, but it only adds a nullable column - 3.0.1 ignores it. Rolling back the code is safe; the column can stay.
15:31  @priya    Probably the database: a query on orders by customer_id is doing a full scan. I can add the index right now in production, a few minutes.
$ incident say "@priya hold the index for now - a CREATE INDEX on that table locks writes"
15:32  @priya    Fair - on a table this size CREATE INDEX without CONCURRENTLY takes a lock. I will prepare it properly for tomorrow.
15:32  @lee      Network team says nothing changed on their side today, for what it is worth.
$ incident decide "roll back checkout to 3.0.1" --why "errors and latency started with 3.0.2 and the rollback is safe; if p99 is not under 800 ms 10 min after it finishes, we look at the database next"
Decision recorded at 15:32 UTC.
$ incident say "Proposal: roll back to 3.0.1. Any strong objections?"
15:34  @dana     No objection from me.

Asked, checked the door, recorded, polled. Total cost: four minutes. PagerDuty's IC guide again: "Making the 'wrong' decision is better than making no decision." A wrong two-way-door decision costs you ten minutes and teaches you something. No decision costs you the whole hour.

When you cannot decide

Sometimes the decision is not yours to make: taking the whole shop offline, spending on emergency capacity, telling customers something legally sensitive. That is hierarchical escalation (37.14): page the person who can decide, and give them a proposal, not a problem - "I recommend we take checkout offline for 20 minutes to stop duplicate charges; I need your yes in 5 minutes".

What you can now do

Why it helps

Every incident asks you to decide before you know: you will not have the root cause for an hour, and users are hurt now. This lesson gives you a way to choose under that pressure that survives the postmortem: generic mitigations first, the rollback bias with its three checks (is it safe, has it been done before, does the timing really match), and a firm no to a one-way door dressed as a quick fix, such as a CREATE INDEX on a big PostgreSQL table without CONCURRENTLY, which blocks writes while it runs.

Interviewers ask "how do you make decisions when you do not know the cause yet?" to see whether you freeze, fixate or gamble big. A decision recorded with its reason, owner, time box and abort condition shows you do none of those, and gives the next IC and the postmortem the "why".

Commands in this lesson

incident

FAQ

Do I need to prove the release caused it before rolling back?

No. If a change went out shortly before the impact started, roll it back first: most incidents follow a change, and a rollback is usually the fastest action with a known result. That default is the rollback bias. Check its limits first: is the rollback safe (a forward schema migration may not run backwards), has the rollback path been used recently, and does the timing really match: errors one minute after a release is a strong hint, last Tuesday's release is not.

What is a generic mitigation?

The SRE Workbook's name for actions that help for many causes without knowing which one it is: roll back the last change, drain traffic away from a bad region or copy, turn off a feature flag, add capacity. The question is not "what is broken?" but "what is the cheapest action that makes users hurt less, that we can undo if we are wrong?". Capture evidence before a restart destroys it.

What are one-way and two-way doors?

Amazon's 2015 shareholder letter split decisions into those that are consequential and hard or impossible to reverse (one-way) and those you can walk back through if they are wrong (two-way). Two-way doors, like rolling back a release, turning a flag off or scaling out, should be decided fast. One-way doors, like restoring a database from backup, deleting data or a schema change under load, deserve a stop, a careful look and often an escalation.

What is the OODA loop?

Colonel John Boyd's loop for acting under uncertainty: observe, orient, decide, act, then observe again. In an incident: observe the impact and the reports, orient on what changed and which hypotheses are still alive, decide one action with an owner, a time box and an abort condition, and let the ops lead act. Keep several hypotheses rather than one theory, and test the cheapest first; fixation is how smart responders lose an hour.

What if the decision is not mine to make?

Taking the whole shop offline, spending on emergency capacity or telling customers something legally sensitive may need someone with more authority. That is hierarchical escalation: page the person who can decide, and bring a proposal, not a problem: "I recommend we take checkout offline for 20 minutes to stop duplicate charges; I need your yes in 5 minutes." Then record their answer on the timeline.

In an interview Mid

How do you make decisions during an incident when you do not know the root cause yet?

Mitigate first, understand later: stop the bleeding, restore service, preserve the evidence. I look for the cheapest action that makes users hurt less and that I can undo - generic mitigations like rolling back, turning a flag off, draining a bad region or scaling out.

If a change shortly preceded the impact, the rollback bias says roll it back without proving it caused the problem, after three checks: is it safe (schema migrations), has the path been used, does the timing really match.

I sort options into two-way doors, decided fast, and one-way doors - a restore that loses data, a locking schema change - where I stop and often escalate. I keep several hypotheses, not one theory, and test the cheapest first.

Every significant decision goes on the record with its reason, an owner, a check ("p99 under 800 ms within 10 minutes") and an abort condition, for example incident decide "roll back checkout to 3.0.1" --why "..." in the lab. Then I propose it and poll for strong objections.

Also asked: When is rolling back not the right first move? · How do you avoid fixating on one theory during an incident? · What do you write down when you make a decision in the middle of an incident?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.