OnCallReady

Lesson 31.8 · Incident Command & Communication · 22 min read

Communication: cadence, audiences and templates

In plain words

You are on a train that stops between stations. Three versions of the next ten minutes: silence, and everyone starts phoning and guessing; "we will be moving in five minutes", and then it is twenty, and nobody believes the next announcement; or "we are held because of a signal problem ahead; the driver is talking to control; next announcement at 8:30", and people relax, because they know what is going on and when they will hear more.

Incident updates are those announcements. Every update has three parts: the impact in the reader's words, what we are doing, and when the next update comes, in UTC. Each audience (the internal channel, the public status page, executives, support) has its own question, so the facts stay the same but the words change. Never promise a fix time; promise the next update, and keep the cadence even when nothing changed.

Why this lesson

Most people who are affected by an incident never see the incident channel. They see a status page, a message from support, an email from their manager. Whether they trust you afterwards depends far more on those messages than on how fast the fix was. An outage with good communication is remembered as "they had a problem and handled it". The same outage with two hours of silence is remembered as "they hid it".

Chapter 0 gave you the shape of one update (impact, status, next). This lesson is the whole communication job: who needs to hear what, how often, in which words, and the mistakes that turn an update into a new problem.

What you need to know already: 0.29 (cadence, the update shape), 37.2 (severity), 37.4 (the comms lead).

Four audiences, four questions

The same incident looks different from each chair:

audience     where they read it          their question                   they do NOT want
-----------  --------------------------  -------------------------------  -----------------------------
internal     the incident channel,       what is going on, who is on it,  to ask "any update?"
             #incidents, engineering     do you need me?
status page  status.shop.lab (public)    is it broken for me? is it you   pods, versions, root causes
             = customers                 or me? when should I try again?  guesses, people's names
executives   email / #exec-updates       how bad for the business, is it  a stack trace; a fake ETA
                                         under control, do you need
                                         anything from me?
support      their own channel / macro   what do I tell a customer right  "investigating" with nothing
                                         now? is there a workaround?      they can say out loud

The facts are the same; the words, the length and the detail are not. An update that tries to serve all four at once serves none: too technical for customers, too vague for engineers, too long for executives.

The shape every update has

Every update, to every audience, carries three things:

impact    who is affected and how        "some customers cannot complete payment at checkout"
action    what we are doing about it     "we have found the cause and are undoing a change"
next      when they will hear from us    "next update at 20:30 UTC"

Impact in the reader's words: the thing they use (checkout, log in, order emails) and what it does to them (fails, is slow, is late). Not "elevated 5xx on checkout-api".

Action is honest about the stage: investigating (we know something is wrong), identified (we know why), monitoring (a fix is out, we are watching), resolved. Say what you do not know: "we do not yet know the cause" is information.

Next is the most neglected and the most important. A promised time turns waiting into a plan: support can tell customers "check back after 20:30", executives stop asking, and the people fixing it are left alone. Always write the zone - incident times are UTC because responders sit in several countries.

Never promise a fix time. "Fixed in 10 minutes" is a guess that becomes a deadline the moment you post it; when it passes, you have a second incident of trust. Promise the next update instead - that is a promise you control. If someone pushes for an ETA, give what you know: "the rollback takes about 6 minutes; if it works, the next update will say so."

The cadence

An update is due on a fixed rhythm whether or not anything changed. "No change: still rolling back, next update at 20:45 UTC" is a complete update. Silence is not neutral: after half an hour without news, people assume the worst and start escalating around you.

How often? Real guidance, for comparison:

For a long incident (a vendor outage running for hours), lengthening the cadence is fine - say so in an update: "We will now update every hour; next update at 23:00 UTC." Changing the rhythm silently is the same as missing it.

The status page

The status page is the customers' view, and it has its own conventions. Atlassian Statuspage, which many companies use, has four incident states and a status per component (each part of the service customers recognise):

incident states:   Investigating -> Identified -> Monitoring -> Resolved
component status:  Operational, Degraded performance, Partial outage, Major outage,
                   Under maintenance

Components carry the impact at a glance: "Checkout and payments: partial outage" is read by more people than any sentence you write. Set them when you post, and set them back to operational when it is over.

In the room, incident update status takes the state and the component:

$ incident update status --state identified --component partial "Some customers cannot complete payment at checkout. We have found the cause and are undoing a recent change; PayPal payments work normally. Next update at 20:30 UTC."
Posted the status update at 20:02 UTC; next update due 20:30 UTC.
20:02  @learner  [status update - identified] Some customers cannot complete payment at checkout. We have found the cause and are undoing a recent change; PayPal payments work normally. Next update at 20:30 UTC.
$ incident statuspage
status.shop.lab - current status (simulator)

  Website                operational
  Checkout and payments  partial outage
  Mobile app             operational

Problems with checkout and payments
  Identified - Some customers cannot complete payment at checkout. We have found the cause and are undoing a recent change; PayPal payments work normally. Next update at 20:30 UTC.
    Posted 2026-09-29 20:02 UTC

(--component partial sets the main component of this scenario - "Checkout and payments" here; --component "Website=degraded" names another one.)

Words that do not belong outside the channel

Inside the response, precise technical words are good. Outside, they are noise at best and alarming at worst. The room's reviewer (a checklist, not a rule of nature) catches the usual ones. A draft the way an engineer writes it at 20:02:

$ incident check status "Checkout pods returning 5xx after 2.9.1 deploy (TaxClient timeouts). Rolling back."
status draft, 11 words (simulator - a reviewer's checklist, not a rule of nature):
  impact (who, how):       yes
  what we are doing:       yes
  next update time:        MISSING
  jargon for this reader:  "pods" -> say "our servers" or say nothing; "5xx" -> "errors" or "failed payments"; "2.9.1" -> "a recent update"; "TaxClient" -> name what the customer sees ("checkout", "card payments")
  blame:                   none
  promised fix time:       none
Fix: it promises no time for the next update; internal jargon for this audience: "pods" (Kubernetes internals), "5xx" (HTTP status codes), "2.9.1" (a version number), "TaxClient" (an internal service or team name).

The same facts, for customers:

$ incident check status "Some customers cannot complete payment at checkout. We have found the cause and are undoing a recent change. Next update at 20:30 UTC."
status draft, 23 words (simulator - a reviewer's checklist, not a rule of nature):
  impact (who, how):       yes
  what we are doing:       yes
  next update time:        in 24 min (20:30)
  jargon for this reader:  none
  blame:                   none
  promised fix time:       none
Ready to post.

What changed: the reader's noun (payment at checkout) instead of ours (pods, 5xx), "a recent change" instead of a version and a class name, and the next update time. Also gone: any guess. "TaxClient timeouts" may be the symptom of something else; once it is on a public page it is your official explanation.

incident check AUDIENCE "draft" runs the same checks the labs use on what you post, without posting anything. The audiences differ: an executive update may say "a release this evening" and name people; a customer update never names people or versions.

Blame does not belong anywhere

One more check runs on every audience, internal included: no person as the cause. "Mo's deploy broke checkout" is a sentence that will be screenshotted, forwarded, and remembered longer than the fix. It is also usually wrong: a deploy pipeline that let a broken release through to a third of customers is the cause worth naming (37.22). Write what happened, not who: "a change released at 19:31".

The four updates for one moment

Here is the same moment of INC-4519 (20:02 UTC, cause known, rollback running) written four ways. Each is complete for its reader.

INTERNAL  Checkout failing for ~31% of customers since 19:33. Cause: 2.9.1's new tax
          service call times out (TaxClient, 2 s). @dana rolling back to 2.9.0, ETA
          ~8 min. @mo on standby. Next update 20:30 UTC.

STATUS    Some customers cannot complete payment at checkout. We have found the cause
          and are undoing a recent change. PayPal payments work normally. Next update
          at 20:30 UTC.

EXEC      About a third of checkouts have failed since 19:33 UTC, so we are losing
          orders. Cause found: a change released tonight. We are undoing it now,
          expected within 10 minutes. Nothing needed from you. Next update 20:30 UTC.

SUPPORT   Tell customers: we know some card payments fail at checkout, it is on our
          side, and it should be fixed within the hour. Workaround: PayPal works. Do not
          promise a time. Status page: status.shop.lab. Next update 20:30 UTC.

Notice the executive one has a number they can act on ("losing orders") and an explicit "nothing needed from you", and the support one is a script, with the workaround first. Mission 37.10 asks you to write these four yourself.

Templates save the first ten minutes

At 3am nobody writes well. Teams keep templates so that the first post goes out in a minute:

[Investigating] We are investigating reports that <what customers see>.
                We will post an update by <HH:MM UTC>.
[Identified]    We have identified the cause of <what customers see> and are
                working on a fix. <workaround if any>. Next update by <HH:MM UTC>.
[Monitoring]    A fix has been applied and <thing> is working again. We are
                monitoring the results. Next update by <HH:MM UTC>.
[Resolved]      This incident has been resolved. Between <start> and <end> UTC,
                <what customers saw>. We apologise for the disruption.

The template is the floor, not the ceiling: always replace the placeholders with what this incident actually does to people.

What you can now do

Why it helps

Most people affected by an incident never see the incident channel; they see a status page, a message from support or an email from their manager. Whether they trust you afterwards depends more on those messages than on how fast the fix was: a well-communicated outage is "they handled it", two hours of silence is "they hid it".

This lesson gives you the skills you will use on every incident and be asked about in interviews: a fixed shape for every update, a cadence you keep or change out loud, the status page states and component statuses customers read at a glance, and a reviewer's eye for jargon ("pods", "5xx", a version number) and blame ("Mo's deploy broke checkout") before you post. incident check lets you run that review on a draft without posting anything.

Commands in this lesson

incident

FAQ

Why should I never give a fix time?

Because "fixed in 10 minutes" is a guess that becomes a deadline the moment you post it, and when it passes you have a second incident, of trust. Promise the next update instead: that is a promise you control. If someone pushes for an ETA, give what you actually know: "the rollback takes about 6 minutes; if it works, the next update will say so."

How often should I post updates?

On a fixed rhythm, whether or not anything changed: "No change: still rolling back, next update at 20:45 UTC" is a complete update. PagerDuty suggests every 20-30 minutes internally during a major incident, and publicly a first post within 5 minutes, then at least every 20 minutes for the first two hours. Atlassian says never more than an hour externally. The lab's policy: SEV1 and SEV2 every 30 minutes, SEV3 every 60.

What do the status page states mean?

Atlassian Statuspage, which many companies use, has four: Investigating (we know something is wrong, post this fast), Identified (we know the cause and are working on a fix), Monitoring (a fix is in place, we are watching) and Resolved (it is over: say when it started and ended, and apologise). Each component also carries a status: Operational, Degraded performance, Partial outage, Major outage or Under maintenance.

Why not explain the technical cause on the status page?

Customers do not know what pods, 5xx or 2.9.1 mean, so it is noise at best and alarming at worst. Worse, an early guess like "TaxClient timeouts" may be the symptom of something else, and once it is on a public page it is your official explanation. Say what the customer sees ("some customers cannot complete payment at checkout") and "a recent change" instead of a version or a class name.

What do executives and support need that customers do not?

Executives need how bad it is for the business ("we are losing orders"), whether it is under control, and an explicit "nothing needed from you", without a stack trace or a fake ETA. Support needs a script they can say out loud, with the workaround first ("PayPal works"), the status page link and the next update time. An executive update may name a release and people; a customer update never names people or versions.

In an interview Mid

How do you communicate during a major incident?

Four audiences, one set of facts, different words: the internal channel (what is going on, who is on it, do you need me), the public status page (is it broken for me, when should I try again), executives (business impact, is it under control, do you need anything from me) and support (what do I tell a customer, is there a workaround).

Every update has the same shape: impact in the reader's words, the action we are taking, and the next update time in UTC. I never promise a fix time; I promise the next update, and post on the cadence even when nothing changed - SEV1 and SEV2 every 30 minutes in our policy. If I lengthen it for a long incident, I say so in an update.

On the status page I move through Investigating, Identified, Monitoring and Resolved, and set each component's status. Outside the channel: no jargon, no guesses at the cause, and no person named as the cause anywhere. A comms lead owns this so the responders are left alone.

Also asked: What would you post on the status page in the first five minutes? · How do you handle an executive who keeps asking for an ETA? · What goes into a good incident update template?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.