OnCallReady

Lesson 31.15 · Incident Command & Communication · 15 min read

Paging, escalation and handoffs

In plain words

Your sink floods at midnight. You could phone the plumber you know personally, who may be asleep after a week of night calls, or on holiday. Or you call the plumbing company's emergency line: it knows who is on duty tonight, and if that person does not pick up within a few minutes, it rings the next one, then the boss. And you tell them "kitchen sink, water on the floor, main tap off", not just "help".

Paging in an incident is the emergency line. Page the team, not your favourite person, and say why. The escalation policy moves to the next level after a timeout, but you can escalate by hand sooner. And when a long incident outlives a shift, the handoff works like nurses changing shift: a fixed briefing on what matters right now, and an explicit "you have it now" before the first one leaves.

Why this lesson

An incident is a team sport played with whoever answers the phone. Three things decide who that is: the page (who gets woken), the escalation (what happens when they do not answer, or cannot help), and the handoff (who takes over when the first person has to stop). Get them wrong and you have the right expert asleep, the wrong one awake, and an IC on hour nine making bad decisions.

What you need to know already: 37.4 (roles; you can only give roles to people in the channel), 37.8 (the update cadence). Chapter 28 covered alerts that page; this lesson is about humans paging humans.

Paging people

Inside an incident you page people for one reason: you need something only they can do or know. Before you page, be able to finish the sentence "I am paging X because...". Then:

Escalation policies

A team's escalation policy says who is paged first and what happens next if nobody acknowledges: level 1 (the primary on-call), level 2 (the secondary), level 3 (often the team's manager). Each level waits for an escalation timeout before the pager moves on by itself. PagerDuty's own on-call guide recommends a 5-minute timeout, and its motto is "Never hesitate to escalate".

In the room:

$ incident oncall
TEAM          LEVEL 1      LEVEL 2      LEVEL 3      ESCALATES AFTER   WHAT
platform      @lee         @casey       @alex        5 min             Kubernetes, CI, the shared platform
database      @priya       @omar        @alex        15 min            the orders and payments databases
checkout      @dana        @mo          @alex        5 min             checkout and cart services
security      @kim         @alex        -            10 min            security incidents, certificates, keys
payments      @nina        @lee         -            10 min            payment providers and the payments API

PEOPLE
  @lee      Lee Novak        SRE, platform team
  @priya    Priya Raman      database on-call (primary)
  @omar     Omar Lindqvist   database on-call (secondary)
  @alex     Alex Chen        engineering manager
  @dana     Dana Okafor      SRE, checkout team
  @kim      Kim Sato         security on-call
  @nina     Nina Petrova     payments engineer (provider integrations)

The database team waits 15 minutes per level - a long time when order pages are failing. You do not have to wait for the policy: incident escalate TEAM pages the next level now. Watch what happens if you do nothing:

$ incident declare --sev 2 "order pages timing out for about one in five customers since 03:05"
Declared INC-4523 SEV2 at 03:12 UTC.
03:12  bot       INC-4523 declared SEV2 by @learner - you are incident commander until you hand it on. Channel #inc-4523-orders. Policy: updates every 30 min (first due 03:42 UTC).
$ incident page database "orders replica lagging 40 min, order pages timing out ~20%, SEV2"
Paged @priya.
03:14  bot       PagerDuty: paged @priya (on call for database): "orders replica lagging 40 min, order pages timing out ~20%, SEV2"
$ incident wait 16
03:15 -> 03:31 UTC
03:29  bot       PagerDuty: priya did not acknowledge within 15 min - escalated to level 2 (@omar) on "database"

Sixteen minutes for an acknowledgement nobody gave. Escalating by hand at minute 5 would have saved ten of them. Priya is not careless: phones die, pagers go to the wrong device, people sleep through vibrations. The policy exists because that happens every week somewhere; escalating early is not an accusation.

Two kinds of escalation

The word means two different things, and both are healthy:

Neither is a sign of failure. The failure is the IC who knows they need help and waits another twenty minutes to be sure.

Follow the sun, and shift changes

Some teams split on-call by time zone so nobody is woken at night: a European team covers its day, then hands to an American team, then (sometimes) to one in Asia - follow the sun. Google's book also sets limits that hold however you split it: at most 25% of an SRE's time on call, and on average no more than two incidents per 12-hour shift, because handling one well (including the postmortem) takes about six hours.

Long incidents outlive shifts. ICS plans in operational periods (typically 12 to 24 hours) for the same reason software teams rotate the IC every few hours: tired people make worse decisions, and they do not notice it themselves. A good IC plans their own replacement before they need it.

Handoffs

A handoff moves a role - most importantly IC - from one person to another while the incident keeps running. It is the moment information gets lost, so it has a fixed form.

ICS calls it transfer of command: a briefing with everything needed to continue, then telling everyone who is in charge now. Google's book adds that the outgoing commander should say it explicitly ("You're now the incident commander, okay?") and "should not leave the call until receiving firm acknowledgment of handoff". PagerDuty's script for the announcement: "Everyone on the call, be advised, at this time I am handing over command to [name]."

The briefing has five parts. Miss one and the new IC finds out the hard way:

impact now        what users see right now (not at the start)
what we know      the current hypothesis, and what has been ruled out
in flight         every open action, and who owns it (@name)
next update       when it is due, and to whom
open decisions    what is waiting to be decided, and the risks to watch

The room's reviewer checks a draft the same way:

$ incident check handoff "Handing over to Casey. Emails are still delayed. Dana and Nina are on it."
handoff draft (simulator - the five things the next IC needs):
  impact right now:                      yes
  what we know (hypothesis, ruled out):  MISSING
  in flight, and who owns it (@name):    yes
  next update due (a time):              MISSING
  open decisions or risks:               MISSING
Add: what we know so far (the current hypothesis, what is ruled out); when the next update is due (a time); the open decisions or risks.
$ incident check handoff "Impact now: order confirmation emails are about 2 hours late for most customers; orders are fine. What we know: the provider throttled us after the 14:00 marketing send, not a bug of ours. In flight: @dana is draining the queue, @nina is asking the provider for 500/min. Next update due 16:55 UTC. Open decision: the 19:00 marketing send - I recommend postponing it."
handoff draft (simulator - the five things the next IC needs):
  impact right now:                      yes
  what we know (hypothesis, ruled out):  yes
  in flight, and who owns it (@name):    yes
  next update due (a time):              yes
  open decisions or risks:               yes
Ready to hand off.

Then the new IC confirms back in their own words, the role changes in the tool, and the channel hears it: "IC is now @casey; next update 16:55 UTC by Casey." Mission 37.16 is a handoff at the end of a shift.

A handoff is also how you take over a room that went wrong (incident 37.17): the outgoing IC may be exhausted, embarrassed, or still on the phone with a vendor. Thank them, take the briefing, and make the change public.

What you can now do

Why it helps

Three things decide who is working on your incident: the page, the escalation and the handoff. Get them wrong and the right expert is asleep, the wrong one is awake, and the IC is on hour nine making bad decisions without noticing. The lab's database team waits 15 minutes per escalation level; escalating by hand at minute 5 saves ten of them, and it is not an accusation.

This lesson also covers the parts interviewers like to probe: functional versus hierarchical escalation, follow-the-sun rotas, Google's limits on on-call load (at most 25% of an SRE's time, on average no more than two incidents per 12-hour shift), and how to hand over command so nothing is lost: impact now, what we know, what is in flight and who owns it, the next update, and the open decisions.

Commands in this lesson

incident

FAQ

Why page the team instead of the person I know is best?

The team's on-call rota knows who is on duty tonight and who is on holiday. Paging your favourite database engineer directly at 3am bypasses it, and they may be the person who was on call last week and is finally sleeping. A team page also follows the escalation policy if nobody answers. Page because you need something only they can do or know, and put the reason in the page.

What is an escalation timeout?

How long one level of an escalation policy waits for an acknowledgement before the pager moves on to the next level by itself. PagerDuty's on-call guide recommends 5 minutes, and its motto is "Never hesitate to escalate". You do not have to wait for the timeout: in the lab incident escalate TEAM pages the next level now. Phones die and people sleep through vibrations; the policy exists because that happens every week.

What is the difference between functional and hierarchical escalation?

Functional escalation brings in different expertise: the ops lead suspects the database, so you page the database on-call, and nobody's authority changes. Hierarchical escalation goes up for a decision or resources you are not allowed to give: emergency spending, taking a region offline, a public statement, waking a vendor's account manager. Both are healthy. The failure is the IC who knows they need help and waits another twenty minutes to be sure.

What is follow the sun?

Splitting on-call by time zone so nobody is woken at night: a European team covers its day, then hands to an American team, sometimes then to one in Asia. Whatever the split, Google's book sets limits: at most 25% of an SRE's time on call, and on average no more than two incidents per 12-hour shift, because handling one well, postmortem included, takes about six hours.

What makes a good handoff?

A fixed briefing with five parts: impact right now, what we know (the hypothesis and what is ruled out), every action in flight with its owner, when the next update is due, and the open decisions and risks. Then an explicit acknowledgement: Google's book says the outgoing commander should not leave until it is firmly acknowledged. Finally tell everyone: "IC is now @casey; next update 16:55 UTC by Casey."

In an interview Mid

How do you hand over incident command at the end of a shift?

A handoff is where information gets lost, so it has a fixed form - ICS calls it transfer of command: a briefing with everything needed to continue, then telling everyone who is in charge.

The briefing has five parts:

  1. impact right now, not at the start;
  2. what we know: the current hypothesis and what is ruled out;
  3. everything in flight, each with its owner by name;
  4. when the next update is due, and to whom;
  5. open decisions and the risks to watch.

The new IC confirms back in their own words, and I do not leave until they have explicitly accepted: "You're now the incident commander, okay?" Then the role changes in the tool and the channel hears it: "IC is now Casey; next update 16:55 UTC by Casey."

I plan my replacement before I need it. Long incidents rotate the IC every few hours, the way ICS plans in operational periods, because tired people make worse decisions and do not notice it themselves.

Also asked: When would you escalate instead of waiting for the on-call to answer? · What is the difference between functional and hierarchical escalation? · How do you keep on-call sustainable for a team?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.