Why severity matters more than it looks
A severity (SEV) is one word that does a lot of work. It decides who is woken up, whether executives hear about it, how often you must post updates, whether the status page changes, whether a postmortem is mandatory, and - for a bank - whether a regulator's clock starts (37.30). Get it wrong one way and the right people sleep through a real outage. Get it wrong the other way and you burn out the people you will need next week.
Chapter 0 gave you a four-level scale and two rules: declare early, and judge by user impact, not by how hard the fix is. This lesson makes it a policy you can apply in ten seconds, under stress, the same way your colleague would.
What you need to know already: 0.29 (severity, declaring), 37.1 (the incident room).
What real companies use
There is no standard; every company writes its own table. Two public ones:
- PagerDuty's Incident Response guide has five levels. SEV-1: "Critical issue that warrants public notification and liaison with executive teams." SEV-2: "Critical system issue actively impacting many customers' ability to use the product." SEV-3: "Stability or minor customer-impacting issues that require immediate attention from service owners." SEV-4: minor issues that need action but do not affect customers' ability to use the product. SEV-5: cosmetic. Anything at SEV-2 or above is a major incident and gets the full incident process.
- Atlassian's incident handbook has three: SEV 1 "a critical incident with very high impact" (a customer-facing service down for all customers, a confidentiality breach, customer data loss), SEV 2 "a major incident with significant impact" (a service unavailable for a subset of customers), SEV 3 "a minor incident with low impact" (a workaround exists, or performance is degraded but usable). SEV 1 and 2 page someone; SEV 3 waits for working hours.
Both describe severity the same way: how badly users are hurt, and how many of them. Neither mentions the cause, the team, or how clever the fix is.
The impact matrix
A table of sentences is hard to apply at 3am. Most teams turn it into a matrix of two questions. This lab's policy:
$ incident matrix
Impact matrix (the lab's severity policy; every company writes its own)
most users a subset a few
(>50%) (5-50%, (<5%, one
a region) customer)
core journey down SEV1 SEV2 SEV3
core journey degraded SEV2 SEV3 SEV3
non-core or internal SEV3 SEV3 SEV4
no user impact yet SEV4 SEV4 SEV4
Any data loss, security breach or money moved wrongly: SEV1, whatever the numbers.
About to hit users within the hour? Judge by the impact it is about to have.
Update cadence: SEV1 every 30 min, SEV2 every 30 min, SEV3 every 60 min, SEV4 none (a ticket).
Rows: how deep. A core journey is something users came to do: log in, search, pay, check out, see their orders. Down means those users cannot complete it at all. Degraded means they can, but slowly, or after retries, or through a painful workaround ("pay by bank transfer through support"). Non-core is everything they can live without for a while: product images, recommendations, a CSV export, the internal admin tool. No user impact yet is a broken thing that redundancy is still hiding (one of three copies crash-looping, all requests fine).
Columns: how wide. Most users (more than half); a subset (5-50%, or everyone in one region or on one platform); a few (under 5%, or a single customer).
Two overrides beat the matrix. Data loss, a security breach, or money moved wrongly (double charges, refunds to the wrong account) are SEV1 even if one customer is affected: the damage cannot be undone by a fix, and lawyers, security and maybe a regulator need to be in the room early. And imminent impact counts: a database disk that fills in 40 minutes, after which checkout stops for everyone, is a SEV1 now, not in 40 minutes. PagerDuty's guide says the same in four words: "Always Assume The Worst" - if you are unsure between two levels, pick the higher one.
Applying it, worked through
Four reports, the way they arrive - half-formed:
"Push notifications are delayed by ~20 min for everyone."
non-core (people still shop without them), most users -> SEV3
"Card payments from one bank fail. That bank is ~3% of our card payments."
core journey down, a few (<5%) -> SEV3
(watch it: if it is the biggest bank in a country, it is a subset -> SEV2)
"Checkout fails for every customer in the EU (~40% of traffic)."
core journey down, a subset (a region) -> SEV2
"The orders database disk is 97% full and grows 1% every 3 minutes."
no impact yet - but in ~9 minutes writes stop and checkout
is down for everyone: imminent, judge by that -> SEV1
Notice what does not appear: whose fault it is, which team owns it, whether you already know the cause, how long the fix will take. A one-line config revert can fix a SEV1; a SEV3 can take a week.
Declaring
With the severity decided, declare in one line: what is broken, for whom, since when. In the room:
$ incident declare --sev 2
incident: declare needs a one-line summary (what is broken, for whom, since when)
$ incident declare --sev 2 "checkout failing for about a third of customers since 19:33"
Declared INC-4519 SEV2 at 19:40 UTC.
19:40 bot INC-4519 declared SEV2 by @learner - you are incident commander until you hand it on. Channel #inc-4519-checkout. Policy: updates every 30 min (first due 20:10 UTC).
19:42 @sam Sam from support - I saw the incident come in. Customers are calling about failed payments. What can I tell them?
Three things happened in one command: you became incident commander (the person who declares owns it until they hand it on), the cadence clock started (the first update is due in 30 minutes, at 20:10), and the incident became visible - support noticed within two minutes. The card now shows it:
$ incident status
INC-4519 checkout failing after release 2.9.1
State: open SEV2 clock 19:42 UTC (+2 min)
Alert: [FIRING] CheckoutErrorBudgetBurn severity=page
Summary: checkout failing for about a third of customers since 19:33
Impact now: 31.2% of checkout requests failing
Roles: IC @learner ops - comms - scribe -
In channel: @learner @sam
Last update: none yet
Next update: due 20:10 UTC (in 28 min)
Decisions: 0 notes: 0
A good summary is short and factual: "checkout failing for about a third of customers since 19:33". Not a guess at the cause ("DB is down") - you will be wrong half the time and the wrong word sticks. Not a mood ("everything is broken!!").
Declare early, downgrade freely
The costs are lopsided. Declaring a SEV2 that turns out to be a SEV4 costs a few people ten minutes and a message: "Downgrading to SEV4: only the internal dashboard was affected." An hour of an undeclared outage costs the first hour of customer trust, a stale status page and a team that started late.
Severity is not a verdict, it is the current best estimate, and it moves both ways. In the room incident sev N changes it, and it insists on a reason:
$ incident sev 1
incident: say why the severity changes: incident sev 1 --why "now failing for all regions"
Every change is announced in the channel and kept in the timeline with its reason, so nobody wonders later why the pager went quiet or loud. Upgrade the moment the impact grows. Downgrade when the impact shrinks - not when the fix is merely in sight (that is what the status page's "monitoring" state is for, 37.8).
Three traps:
- Severity by seniority. "The CTO noticed, make it SEV1." Who noticed changes who you update, not how bad it is for users.
- Severity by effort. "It is a one-line fix, SEV3." The fix does not reduce the impact until it is out.
- Debating it. PagerDuty's IC guide: do not discuss severity during the call. Pick one from the policy, move on, adjust later.
What you can now do
- place an incident on the impact matrix by depth (down / degraded / non-core / none yet) and breadth (most / a subset / a few), and apply the two overrides
- compare the lab's policy with PagerDuty's five levels and Atlassian's three
- declare with a one-line summary (what, for whom, since when) and change severity with a reason
- spot severity by seniority, by effort, and the severity debate