Why this lesson
An incident is a team sport played with whoever answers the phone. Three things decide who that is: the page (who gets woken), the escalation (what happens when they do not answer, or cannot help), and the handoff (who takes over when the first person has to stop). Get them wrong and you have the right expert asleep, the wrong one awake, and an IC on hour nine making bad decisions.
What you need to know already: 37.4 (roles; you can only give roles to people in the channel), 37.8 (the update cadence). Chapter 28 covered alerts that page; this lesson is about humans paging humans.
Paging people
Inside an incident you page people for one reason: you need something only they can do or know. Before you page, be able to finish the sentence "I am paging X because...". Then:
- Page the team, not the person. The team's on-call rota knows who is on duty tonight and who is on holiday. Paging your favourite database engineer directly at 3am bypasses the rota - and they may be the one who was on call last week and is finally sleeping.
- Say why in the page. "orders replica 40 min behind, order pages timing out for ~20%, SEV2 declared, need the database on-call" lets them start thinking on the way to the laptop. "help" does not.
- Page enough, not everyone. Each extra person paged is someone woken, someone to brief, and one more voice in the channel. Paging security for a database lag, or your manager for a SEV3, is noise.
- Know the clock. Google's SRE book gives typical expected response times: 5 minutes for user-facing or time-critical services, 30 minutes for less time-sensitive ones. That is the window before you should start worrying.
Escalation policies
A team's escalation policy says who is paged first and what happens next if nobody acknowledges: level 1 (the primary on-call), level 2 (the secondary), level 3 (often the team's manager). Each level waits for an escalation timeout before the pager moves on by itself. PagerDuty's own on-call guide recommends a 5-minute timeout, and its motto is "Never hesitate to escalate".
In the room:
$ incident oncall
TEAM LEVEL 1 LEVEL 2 LEVEL 3 ESCALATES AFTER WHAT
platform @lee @casey @alex 5 min Kubernetes, CI, the shared platform
database @priya @omar @alex 15 min the orders and payments databases
checkout @dana @mo @alex 5 min checkout and cart services
security @kim @alex - 10 min security incidents, certificates, keys
payments @nina @lee - 10 min payment providers and the payments API
PEOPLE
@lee Lee Novak SRE, platform team
@priya Priya Raman database on-call (primary)
@omar Omar Lindqvist database on-call (secondary)
@alex Alex Chen engineering manager
@dana Dana Okafor SRE, checkout team
@kim Kim Sato security on-call
@nina Nina Petrova payments engineer (provider integrations)
The database team waits 15 minutes per level - a long time when order pages are failing. You do not have to wait for the policy: incident escalate TEAM pages the next level now. Watch what happens if you do nothing:
$ incident declare --sev 2 "order pages timing out for about one in five customers since 03:05"
Declared INC-4523 SEV2 at 03:12 UTC.
03:12 bot INC-4523 declared SEV2 by @learner - you are incident commander until you hand it on. Channel #inc-4523-orders. Policy: updates every 30 min (first due 03:42 UTC).
$ incident page database "orders replica lagging 40 min, order pages timing out ~20%, SEV2"
Paged @priya.
03:14 bot PagerDuty: paged @priya (on call for database): "orders replica lagging 40 min, order pages timing out ~20%, SEV2"
$ incident wait 16
03:15 -> 03:31 UTC
03:29 bot PagerDuty: priya did not acknowledge within 15 min - escalated to level 2 (@omar) on "database"
Sixteen minutes for an acknowledgement nobody gave. Escalating by hand at minute 5 would have saved ten of them. Priya is not careless: phones die, pagers go to the wrong device, people sleep through vibrations. The policy exists because that happens every week somewhere; escalating early is not an accusation.
Two kinds of escalation
The word means two different things, and both are healthy:
- Functional escalation: you need different expertise. The ops lead suspects the database; you page the database on-call. Nobody's authority changes.
- Hierarchical escalation: you need a decision or resources you are not allowed to give: spend money on emergency capacity, take a whole region offline, make a public statement, wake a vendor's account manager. You escalate to a manager or an executive.
Neither is a sign of failure. The failure is the IC who knows they need help and waits another twenty minutes to be sure.
Follow the sun, and shift changes
Some teams split on-call by time zone so nobody is woken at night: a European team covers its day, then hands to an American team, then (sometimes) to one in Asia - follow the sun. Google's book also sets limits that hold however you split it: at most 25% of an SRE's time on call, and on average no more than two incidents per 12-hour shift, because handling one well (including the postmortem) takes about six hours.
Long incidents outlive shifts. ICS plans in operational periods (typically 12 to 24 hours) for the same reason software teams rotate the IC every few hours: tired people make worse decisions, and they do not notice it themselves. A good IC plans their own replacement before they need it.
Handoffs
A handoff moves a role - most importantly IC - from one person to another while the incident keeps running. It is the moment information gets lost, so it has a fixed form.
ICS calls it transfer of command: a briefing with everything needed to continue, then telling everyone who is in charge now. Google's book adds that the outgoing commander should say it explicitly ("You're now the incident commander, okay?") and "should not leave the call until receiving firm acknowledgment of handoff". PagerDuty's script for the announcement: "Everyone on the call, be advised, at this time I am handing over command to [name]."
The briefing has five parts. Miss one and the new IC finds out the hard way:
impact now what users see right now (not at the start)
what we know the current hypothesis, and what has been ruled out
in flight every open action, and who owns it (@name)
next update when it is due, and to whom
open decisions what is waiting to be decided, and the risks to watch
The room's reviewer checks a draft the same way:
$ incident check handoff "Handing over to Casey. Emails are still delayed. Dana and Nina are on it."
handoff draft (simulator - the five things the next IC needs):
impact right now: yes
what we know (hypothesis, ruled out): MISSING
in flight, and who owns it (@name): yes
next update due (a time): MISSING
open decisions or risks: MISSING
Add: what we know so far (the current hypothesis, what is ruled out); when the next update is due (a time); the open decisions or risks.
$ incident check handoff "Impact now: order confirmation emails are about 2 hours late for most customers; orders are fine. What we know: the provider throttled us after the 14:00 marketing send, not a bug of ours. In flight: @dana is draining the queue, @nina is asking the provider for 500/min. Next update due 16:55 UTC. Open decision: the 19:00 marketing send - I recommend postponing it."
handoff draft (simulator - the five things the next IC needs):
impact right now: yes
what we know (hypothesis, ruled out): yes
in flight, and who owns it (@name): yes
next update due (a time): yes
open decisions or risks: yes
Ready to hand off.
Then the new IC confirms back in their own words, the role changes in the tool, and the channel hears it: "IC is now @casey; next update 16:55 UTC by Casey." Mission 37.16 is a handoff at the end of a shift.
A handoff is also how you take over a room that went wrong (incident 37.17): the outgoing IC may be exhausted, embarrassed, or still on the phone with a vendor. Thank them, take the briefing, and make the change public.
What you can now do
- page the right team with a reason, and know the expected response times
- read an escalation policy, and escalate by hand instead of waiting out a long timeout
- tell functional from hierarchical escalation
- explain follow-the-sun, on-call load limits and why long incidents rotate the IC
- hand off command with the five parts and get an explicit acknowledgement