OnCallReady

Lesson 31.2 · Incident Command & Communication · 16 min read

Severity levels: the impact matrix, and declaring early

In plain words

In a hospital emergency room, a triage nurse decides who is seen first. They do not ask who shouts loudest, who is famous, or how easy the treatment will be. They look at two things: how badly hurt the person is, and how urgent it is. A patient who looks fine but is about to collapse goes ahead of a loud sprained ankle.

Severity works the same way. The lab's impact matrix asks how deep the harm is (a core journey down, degraded, something non-core, or no user impact yet) and how wide (most users, a subset or a region, a few). Two overrides beat the table: data loss, a security breach or money moved wrongly is SEV1 whatever the numbers, and impact about to happen counts as if it were happening. Then you declare in one line: what is broken, for whom, since when.

Why severity matters more than it looks

A severity (SEV) is one word that does a lot of work. It decides who is woken up, whether executives hear about it, how often you must post updates, whether the status page changes, whether a postmortem is mandatory, and - for a bank - whether a regulator's clock starts (37.30). Get it wrong one way and the right people sleep through a real outage. Get it wrong the other way and you burn out the people you will need next week.

Chapter 0 gave you a four-level scale and two rules: declare early, and judge by user impact, not by how hard the fix is. This lesson makes it a policy you can apply in ten seconds, under stress, the same way your colleague would.

What you need to know already: 0.29 (severity, declaring), 37.1 (the incident room).

What real companies use

There is no standard; every company writes its own table. Two public ones:

Both describe severity the same way: how badly users are hurt, and how many of them. Neither mentions the cause, the team, or how clever the fix is.

The impact matrix

A table of sentences is hard to apply at 3am. Most teams turn it into a matrix of two questions. This lab's policy:

$ incident matrix
Impact matrix (the lab's severity policy; every company writes its own)

                            most users  a subset    a few
                            (>50%)      (5-50%,     (<5%, one
                                        a region)   customer)
  core journey down         SEV1        SEV2        SEV3
  core journey degraded     SEV2        SEV3        SEV3
  non-core or internal      SEV3        SEV3        SEV4
  no user impact yet        SEV4        SEV4        SEV4

  Any data loss, security breach or money moved wrongly: SEV1, whatever the numbers.
  About to hit users within the hour? Judge by the impact it is about to have.

Update cadence: SEV1 every 30 min, SEV2 every 30 min, SEV3 every 60 min, SEV4 none (a ticket).

Rows: how deep. A core journey is something users came to do: log in, search, pay, check out, see their orders. Down means those users cannot complete it at all. Degraded means they can, but slowly, or after retries, or through a painful workaround ("pay by bank transfer through support"). Non-core is everything they can live without for a while: product images, recommendations, a CSV export, the internal admin tool. No user impact yet is a broken thing that redundancy is still hiding (one of three copies crash-looping, all requests fine).

Columns: how wide. Most users (more than half); a subset (5-50%, or everyone in one region or on one platform); a few (under 5%, or a single customer).

Two overrides beat the matrix. Data loss, a security breach, or money moved wrongly (double charges, refunds to the wrong account) are SEV1 even if one customer is affected: the damage cannot be undone by a fix, and lawyers, security and maybe a regulator need to be in the room early. And imminent impact counts: a database disk that fills in 40 minutes, after which checkout stops for everyone, is a SEV1 now, not in 40 minutes. PagerDuty's guide says the same in four words: "Always Assume The Worst" - if you are unsure between two levels, pick the higher one.

Applying it, worked through

Four reports, the way they arrive - half-formed:

"Push notifications are delayed by ~20 min for everyone."
    non-core (people still shop without them), most users     -> SEV3

"Card payments from one bank fail. That bank is ~3% of our card payments."
    core journey down, a few (<5%)                             -> SEV3
    (watch it: if it is the biggest bank in a country, it is a subset -> SEV2)

"Checkout fails for every customer in the EU (~40% of traffic)."
    core journey down, a subset (a region)                     -> SEV2

"The orders database disk is 97% full and grows 1% every 3 minutes."
    no impact yet - but in ~9 minutes writes stop and checkout
    is down for everyone: imminent, judge by that              -> SEV1

Notice what does not appear: whose fault it is, which team owns it, whether you already know the cause, how long the fix will take. A one-line config revert can fix a SEV1; a SEV3 can take a week.

Declaring

With the severity decided, declare in one line: what is broken, for whom, since when. In the room:

$ incident declare --sev 2
incident: declare needs a one-line summary (what is broken, for whom, since when)
$ incident declare --sev 2 "checkout failing for about a third of customers since 19:33"
Declared INC-4519 SEV2 at 19:40 UTC.
19:40  bot       INC-4519 declared SEV2 by @learner - you are incident commander until you hand it on. Channel #inc-4519-checkout. Policy: updates every 30 min (first due 20:10 UTC).
19:42  @sam      Sam from support - I saw the incident come in. Customers are calling about failed payments. What can I tell them?

Three things happened in one command: you became incident commander (the person who declares owns it until they hand it on), the cadence clock started (the first update is due in 30 minutes, at 20:10), and the incident became visible - support noticed within two minutes. The card now shows it:

$ incident status
INC-4519  checkout failing after release 2.9.1
State:        open  SEV2   clock 19:42 UTC (+2 min)
Alert:        [FIRING] CheckoutErrorBudgetBurn severity=page
Summary:      checkout failing for about a third of customers since 19:33
Impact now:   31.2% of checkout requests failing
Roles:        IC @learner   ops -   comms -   scribe -
In channel:   @learner @sam
Last update:  none yet
Next update:  due 20:10 UTC (in 28 min)
Decisions:    0   notes: 0

A good summary is short and factual: "checkout failing for about a third of customers since 19:33". Not a guess at the cause ("DB is down") - you will be wrong half the time and the wrong word sticks. Not a mood ("everything is broken!!").

Declare early, downgrade freely

The costs are lopsided. Declaring a SEV2 that turns out to be a SEV4 costs a few people ten minutes and a message: "Downgrading to SEV4: only the internal dashboard was affected." An hour of an undeclared outage costs the first hour of customer trust, a stale status page and a team that started late.

Severity is not a verdict, it is the current best estimate, and it moves both ways. In the room incident sev N changes it, and it insists on a reason:

$ incident sev 1
incident: say why the severity changes: incident sev 1 --why "now failing for all regions"

Every change is announced in the channel and kept in the timeline with its reason, so nobody wonders later why the pager went quiet or loud. Upgrade the moment the impact grows. Downgrade when the impact shrinks - not when the fix is merely in sight (that is what the status page's "monitoring" state is for, 37.8).

Three traps:

What you can now do

Why it helps

A severity is one word that decides who is woken, whether executives hear about it, how often you post updates, whether the status page changes, whether a postmortem is mandatory and, at a bank, whether a regulator's clock starts. Too low and the right people sleep through an outage; too high and you burn out the people you need next week.

A matrix lets you decide in ten seconds, at 3am, the same way your colleague would, and it keeps out the three traps: severity by seniority ("the CTO noticed"), severity by effort ("it is a one-line fix") and debating it on the call. Interviewers ask "how do you decide severity?" to see whether you judge by user impact; comparing the lab's policy with PagerDuty's five levels and Atlassian's three shows you know there is no single standard.

Commands in this lesson

incident

FAQ

Is there a standard severity scale?

No, every company writes its own. PagerDuty's public guide has five levels, SEV-1 (warrants public notification and liaison with executives) to SEV-5 (cosmetic), and anything at SEV-2 or above is a major incident with the full process. Atlassian's handbook has three: SEV 1 critical, SEV 2 major, SEV 3 minor; SEV 1 and 2 page someone, SEV 3 waits for working hours. Both judge by how badly users are hurt and how many.

What counts as a core journey?

Something users came to do: log in, search, pay, check out, see their orders. Down means they cannot complete it at all. Degraded means they can, but slowly, after retries, or through a painful workaround such as paying by bank transfer through support. Non-core is what they can live without for a while: product images, recommendations, a CSV export, the internal admin tool.

Why is a nearly full database disk a SEV1 when nobody is affected yet?

Because imminent impact counts. In the worked example the orders database disk is 97% full and grows 1% every 3 minutes: in about 9 minutes writes stop and checkout is down for everyone. You judge by the impact it is about to have, so it is a SEV1 now, not in 9 minutes. PagerDuty's guide says "Always Assume The Worst": unsure between two levels, pick the higher one.

What if I pick the wrong severity?

Change it. Severity is the current best estimate, not a verdict, and it moves both ways. Declaring a SEV2 that turns out to be a SEV4 costs a few people ten minutes and a message; an undeclared hour costs customer trust. In the room incident sev N changes it and insists on a reason with --why, so the channel and the timeline say why the pager went loud or quiet.

Should severity depend on who noticed or how hard the fix is?

Neither. Who noticed changes who you update, not how bad it is for users: the CTO spotting it does not make it a SEV1. The fix does not reduce the impact until it is out, so "it is a one-line fix, SEV3" is wrong too: a one-line revert can fix a SEV1 and a SEV3 can take a week. And do not debate it on the call: pick one from the policy and adjust later.

In an interview Junior

How do you decide the severity of an incident?

By user impact, never by the cause, the team, who noticed or how hard the fix is. I use the company's matrix with two questions:

Core journey down for most users is a SEV1; degraded for one region is a SEV3 in the lab's policy. Two overrides beat the table: data loss, a security breach or money moved wrongly is SEV1 even for one customer, and imminent impact counts - a disk that fills in 9 minutes is a SEV1 now. Unsure between two levels, pick the higher one.

Then I declare in one line, what is broken, for whom, since when - incident declare --sev 2 "checkout failing for about a third of customers since 19:33" in the lab - and change the severity with a stated reason as the impact moves.

Also asked: What is the difference between a SEV1 and a SEV3? · Why declare an incident before you know the cause? · When would you downgrade an incident's severity?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.