$ incident check status "Checkout pods returning 5xx after 2.9.1 deploy (TaxClient timeouts). Rolling back."
status draft, 11 words (simulator - a reviewer's checklist, not a rule of nature):
impact (who, how): yes
what we are doing: yes
next update time: MISSING
jargon for this reader: "pods" -> say "our servers" or say nothing; "5xx" -> "errors" or "failed payments"; "2.9.1" -> "a recent update"; "TaxClient" -> name what the customer sees ("checkout", "card payments")
blame: none
promised fix time: noneThat is a real first draft, written at 20:02 during a checkout outage by an engineer who knew exactly what was wrong (a slow new dependency, the kind of thing a p99 latency graph shows first). It is accurate. It is also useless on a public status page: no customer knows what a pod or a 5xx is, and nothing in it says when they should come back.
The three parts every update has
Whatever the audience, an incident update carries three things:
- Impact - who is affected and how, in the reader's words: "some customers cannot complete payment at checkout", not "elevated 5xx on checkout-api".
- Action - what you are doing, honest about the stage: investigating (something is wrong), identified (we know why), monitoring (a fix is in, we are watching), resolved.
- Next - when the next update comes, with a time zone: "next update at 20:30 UTC".
The third is the one people forget, and the one that matters most. A promised time turns waiting into a plan: support can tell customers to check back after 20:30, executives stop asking, and the people fixing it are left alone.
$ incident check status "Some customers cannot complete payment at checkout. We have found the cause and are undoing a recent change. Next update at 20:30 UTC."
status draft, 23 words (simulator - a reviewer's checklist, not a rule of nature):
impact (who, how): yes
what we are doing: yes
next update time: in 28 min (20:30)
jargon for this reader: none
blame: none
promised fix time: none
Ready to post.Never promise a fix time
"It will be fixed in 10 minutes" is a guess that becomes a deadline the moment it is posted. When it passes, you have a second incident, this one about trust. Promise the next update instead: that is a promise you control. The x509 Monday outage is a good example: the certificate renewal takes about 25 minutes, but nobody can promise when the cluster has scaled back up. If someone pushes for an ETA, give what you know ("the rollback takes about six minutes; the next update will say whether it worked").
A cadence, even when nothing changed
Updates go out on a fixed rhythm whether or not there is news. "No change: we are still rolling back. Next update at 20:45 UTC." is a complete update. PagerDuty's public incident response guide suggests every 20 to 30 minutes during a major incident, with the first public post within five minutes of starting the response; Atlassian's handbook says never more than an hour externally without an update. Silence is not neutral: after half an hour without news, people assume the worst and escalate around you - and for a customer staring at a 502 Bad Gateway, the status page is the only sign anyone knows.
The same moment, four audiences
The facts stay the same; the words do not.
INTERNAL Checkout failing for ~31% of customers since 19:33. Cause: 2.9.1's new tax
service call times out (2 s). @dana rolling back to 2.9.0, ~8 min. Next
update 20:30 UTC.
STATUS Some customers cannot complete payment at checkout. We have found the cause
and are undoing a recent change. PayPal payments work normally. Next update
at 20:30 UTC.
EXEC About a third of checkouts have failed since 19:33 UTC, so we are losing
orders. Cause found: a change released tonight; we are undoing it now.
Nothing needed from you. Next update 20:30 UTC.
SUPPORT Tell customers: some card payments fail at checkout, it is on our side and
we are fixing it. Workaround: PayPal works. Do not promise a time. Next
update 20:30 UTC.Inside the response, versions and names help. For customers, the workaround is the most useful sentence. For executives, business impact and an explicit "nothing needed from you" stop the follow-up questions. For support, a script they can read aloud. All four carry the same next update time, or three of them will be wrong.
Status page states
Most status pages follow Atlassian Statuspage's model: Investigating, Identified, Monitoring, Resolved, plus a status per component (operational, degraded performance, partial outage, major outage). Post Investigating before you know the cause. Post Resolved only after the fix has held for a while, with when it started and ended, what customers saw, and a plain apology.
Practise it
OnCallReady's chapter 37 has an incident room with a channel, a pager and a status page, on an incident clock: responders reply to what you post, support asks for news when a promised update is late, and the labs check every update you send for impact, action, next time, jargon and blame. The lesson "Communication: cadence, audiences and templates" walks through all of it.