Why this lesson
Most people who are affected by an incident never see the incident channel. They see a status page, a message from support, an email from their manager. Whether they trust you afterwards depends far more on those messages than on how fast the fix was. An outage with good communication is remembered as "they had a problem and handled it". The same outage with two hours of silence is remembered as "they hid it".
Chapter 0 gave you the shape of one update (impact, status, next). This lesson is the whole communication job: who needs to hear what, how often, in which words, and the mistakes that turn an update into a new problem.
What you need to know already: 0.29 (cadence, the update shape), 37.2 (severity), 37.4 (the comms lead).
Four audiences, four questions
The same incident looks different from each chair:
audience where they read it their question they do NOT want
----------- -------------------------- ------------------------------- -----------------------------
internal the incident channel, what is going on, who is on it, to ask "any update?"
#incidents, engineering do you need me?
status page status.shop.lab (public) is it broken for me? is it you pods, versions, root causes
= customers or me? when should I try again? guesses, people's names
executives email / #exec-updates how bad for the business, is it a stack trace; a fake ETA
under control, do you need
anything from me?
support their own channel / macro what do I tell a customer right "investigating" with nothing
now? is there a workaround? they can say out loud
The facts are the same; the words, the length and the detail are not. An update that tries to serve all four at once serves none: too technical for customers, too vague for engineers, too long for executives.
The shape every update has
Every update, to every audience, carries three things:
impact who is affected and how "some customers cannot complete payment at checkout"
action what we are doing about it "we have found the cause and are undoing a change"
next when they will hear from us "next update at 20:30 UTC"
Impact in the reader's words: the thing they use (checkout, log in, order emails) and what it does to them (fails, is slow, is late). Not "elevated 5xx on checkout-api".
Action is honest about the stage: investigating (we know something is wrong), identified (we know why), monitoring (a fix is out, we are watching), resolved. Say what you do not know: "we do not yet know the cause" is information.
Next is the most neglected and the most important. A promised time turns waiting into a plan: support can tell customers "check back after 20:30", executives stop asking, and the people fixing it are left alone. Always write the zone - incident times are UTC because responders sit in several countries.
Never promise a fix time. "Fixed in 10 minutes" is a guess that becomes a deadline the moment you post it; when it passes, you have a second incident of trust. Promise the next update instead - that is a promise you control. If someone pushes for an ETA, give what you know: "the rollback takes about 6 minutes; if it works, the next update will say so."
The cadence
An update is due on a fixed rhythm whether or not anything changed. "No change: still rolling back, next update at 20:45 UTC" is a complete update. Silence is not neutral: after half an hour without news, people assume the worst and start escalating around you.
How often? Real guidance, for comparison:
- PagerDuty's guide: status updates every 20-30 minutes during a major incident internally; publicly, the first post within 5 minutes of starting the incident call, then at least every 20 minutes during the first two hours.
- Atlassian's handbook: externally, never go more than one hour without an update, and always say when the next one will be.
- This lab's policy (
incident matrix): SEV1 and SEV2 every 30 minutes, SEV3 every 60, SEV4 none. The room tracks it: the card shows when the next update is due, and the time you promise in the text counts if it is sooner.
For a long incident (a vendor outage running for hours), lengthening the cadence is fine - say so in an update: "We will now update every hour; next update at 23:00 UTC." Changing the rhythm silently is the same as missing it.
The status page
The status page is the customers' view, and it has its own conventions. Atlassian Statuspage, which many companies use, has four incident states and a status per component (each part of the service customers recognise):
incident states: Investigating -> Identified -> Monitoring -> Resolved
component status: Operational, Degraded performance, Partial outage, Major outage,
Under maintenance
- Investigating: we know something is wrong. Post this fast - before you know why.
- Identified: we know the cause and are working on the fix.
- Monitoring: a fix is in place; we are watching to be sure.
- Resolved: it is over. Say when it started and ended, and apologise plainly.
Components carry the impact at a glance: "Checkout and payments: partial outage" is read by more people than any sentence you write. Set them when you post, and set them back to operational when it is over.
In the room, incident update status takes the state and the component:
$ incident update status --state identified --component partial "Some customers cannot complete payment at checkout. We have found the cause and are undoing a recent change; PayPal payments work normally. Next update at 20:30 UTC."
Posted the status update at 20:02 UTC; next update due 20:30 UTC.
20:02 @learner [status update - identified] Some customers cannot complete payment at checkout. We have found the cause and are undoing a recent change; PayPal payments work normally. Next update at 20:30 UTC.
$ incident statuspage
status.shop.lab - current status (simulator)
Website operational
Checkout and payments partial outage
Mobile app operational
Problems with checkout and payments
Identified - Some customers cannot complete payment at checkout. We have found the cause and are undoing a recent change; PayPal payments work normally. Next update at 20:30 UTC.
Posted 2026-09-29 20:02 UTC
(--component partial sets the main component of this scenario - "Checkout and payments" here; --component "Website=degraded" names another one.)
Words that do not belong outside the channel
Inside the response, precise technical words are good. Outside, they are noise at best and alarming at worst. The room's reviewer (a checklist, not a rule of nature) catches the usual ones. A draft the way an engineer writes it at 20:02:
$ incident check status "Checkout pods returning 5xx after 2.9.1 deploy (TaxClient timeouts). Rolling back."
status draft, 11 words (simulator - a reviewer's checklist, not a rule of nature):
impact (who, how): yes
what we are doing: yes
next update time: MISSING
jargon for this reader: "pods" -> say "our servers" or say nothing; "5xx" -> "errors" or "failed payments"; "2.9.1" -> "a recent update"; "TaxClient" -> name what the customer sees ("checkout", "card payments")
blame: none
promised fix time: none
Fix: it promises no time for the next update; internal jargon for this audience: "pods" (Kubernetes internals), "5xx" (HTTP status codes), "2.9.1" (a version number), "TaxClient" (an internal service or team name).
The same facts, for customers:
$ incident check status "Some customers cannot complete payment at checkout. We have found the cause and are undoing a recent change. Next update at 20:30 UTC."
status draft, 23 words (simulator - a reviewer's checklist, not a rule of nature):
impact (who, how): yes
what we are doing: yes
next update time: in 24 min (20:30)
jargon for this reader: none
blame: none
promised fix time: none
Ready to post.
What changed: the reader's noun (payment at checkout) instead of ours (pods, 5xx), "a recent change" instead of a version and a class name, and the next update time. Also gone: any guess. "TaxClient timeouts" may be the symptom of something else; once it is on a public page it is your official explanation.
incident check AUDIENCE "draft" runs the same checks the labs use on what you post, without posting anything. The audiences differ: an executive update may say "a release this evening" and name people; a customer update never names people or versions.
Blame does not belong anywhere
One more check runs on every audience, internal included: no person as the cause. "Mo's deploy broke checkout" is a sentence that will be screenshotted, forwarded, and remembered longer than the fix. It is also usually wrong: a deploy pipeline that let a broken release through to a third of customers is the cause worth naming (37.22). Write what happened, not who: "a change released at 19:31".
The four updates for one moment
Here is the same moment of INC-4519 (20:02 UTC, cause known, rollback running) written four ways. Each is complete for its reader.
INTERNAL Checkout failing for ~31% of customers since 19:33. Cause: 2.9.1's new tax
service call times out (TaxClient, 2 s). @dana rolling back to 2.9.0, ETA
~8 min. @mo on standby. Next update 20:30 UTC.
STATUS Some customers cannot complete payment at checkout. We have found the cause
and are undoing a recent change. PayPal payments work normally. Next update
at 20:30 UTC.
EXEC About a third of checkouts have failed since 19:33 UTC, so we are losing
orders. Cause found: a change released tonight. We are undoing it now,
expected within 10 minutes. Nothing needed from you. Next update 20:30 UTC.
SUPPORT Tell customers: we know some card payments fail at checkout, it is on our
side, and it should be fixed within the hour. Workaround: PayPal works. Do not
promise a time. Status page: status.shop.lab. Next update 20:30 UTC.
Notice the executive one has a number they can act on ("losing orders") and an explicit "nothing needed from you", and the support one is a script, with the workaround first. Mission 37.10 asks you to write these four yourself.
Templates save the first ten minutes
At 3am nobody writes well. Teams keep templates so that the first post goes out in a minute:
[Investigating] We are investigating reports that <what customers see>.
We will post an update by <HH:MM UTC>.
[Identified] We have identified the cause of <what customers see> and are
working on a fix. <workaround if any>. Next update by <HH:MM UTC>.
[Monitoring] A fix has been applied and <thing> is working again. We are
monitoring the results. Next update by <HH:MM UTC>.
[Resolved] This incident has been resolved. Between <start> and <end> UTC,
<what customers saw>. We apologise for the disruption.
The template is the floor, not the ceiling: always replace the placeholders with what this incident actually does to people.
What you can now do
- name the four audiences and the question each brings
- write an update with impact, action and next time, in the reader's words, with UTC times and no fix-time promise
- keep (and explicitly change) a cadence; compare it with PagerDuty's and Atlassian's
- drive a status page through investigating / identified / monitoring / resolved with component statuses
- spot jargon and blame before you post, with
incident check