Incident Command & Communication: interview questions
The question you are most likely to get for each topic, a model answer, and what else comes up. From chapter 31 of the course.
Walk me through how you would run a major incident, from the page to the postmortem. Mid
- Declare early. Severity from user impact (how deep, how wide), not from the cause; a one-line summary: what is broken, for whom, since when. I am IC until I hand it on.
- Roles. Page the team, not a person. One ops lead is the only route to production (unity of command); then a comms lead and a scribe. I do not debug.
- Tasks with who, what and when, and CAN reports back: condition, actions, needs.
- Updates on the cadence to every audience: impact, what we are doing, the next update time in UTC. Never a fix-time promise.
- Mitigate first. If a change preceded the impact, roll it back. Two-way doors fast; one-way doors stop and escalate. Each decision recorded with its reason and an abort condition.
- Escalate by hand rather than wait out a timeout; hand off with a briefing and an explicit acknowledgement.
- Resolve when the impact has ended, close every audience, then a blameless postmortem: contributing factors, and action items with an owner, a date and a way to tell they are done.
Also asked: How do you decide the severity of an incident? · What does an incident commander do if they do not fix anything? · How would you measure how well your team handles incidents?
Why should the incident commander not debug? Junior
Debugging needs deep focus on one thing; commanding needs shallow attention on everything, and the same brain cannot do both. The moment the IC opens a terminal, "what is the impact now", "who is doing what" and "when is the next update" have no owner. That is hero mode again, just with a title.
PagerDuty's guide says the IC is "NOT a resolver" and should not be checking graphs or investigating logs. Google's book says the operations team should be the only group modifying the system during an incident - unity of command: one IC, and only the ops lead's people touch production. The IC's job is the 3Cs: coordinate, communicate, control.
On a tiny team I would wear both hats, but say which one I have on and take the IC hat back every few minutes. As soon as a second person arrives, IC and ops are the first roles to split.
Also asked: Have you ever been incident commander, and what did you do? · When would you declare an incident instead of just fixing it quietly? · What goes wrong when several engineers change production at the same time?
Learn it: 31.1 Why incident command: from wildfires to the IC who does not debug
How do you decide the severity of an incident? Junior
By user impact, never by the cause, the team, who noticed or how hard the fix is. I use the company's matrix with two questions:
- How deep: a core journey (log in, pay, check out) down, degraded (slow, needs retries), something non-core, or no user impact yet.
- How wide: most users, a subset or a region, or a few.
Core journey down for most users is a SEV1; degraded for one region is a SEV3 in the lab's policy. Two overrides beat the table: data loss, a security breach or money moved wrongly is SEV1 even for one customer, and imminent impact counts - a disk that fills in 9 minutes is a SEV1 now. Unsure between two levels, pick the higher one.
Then I declare in one line, what is broken, for whom, since when - incident declare --sev 2 "checkout failing for about a third of customers since 19:33" in the lab - and change the severity with a stated reason as the impact moves.
Also asked: What is the difference between a SEV1 and a SEV3? · Why declare an incident before you know the cause? · When would you downgrade an incident's severity?
Learn it: 31.2 Severity levels: the impact matrix, and declaring early
What does an incident commander actually do if they do not fix anything? Junior
Coordinate, communicate, control. Concretely:
- set and change the severity, and declare when it is over;
- hand out roles by name: one ops lead, the only hands on production (unity of command), then a comms lead and a scribe; until a role is filled, it is mine;
- give tasks with who, what and when ("@dana, read the checkout errors, report back in 5 minutes"), and ask for a CAN report: condition, actions, needs;
- stop uncoordinated changes ("nobody touches prod unless I assign it");
- keep the update cadence for every audience;
- make decisions, proposing and polling for strong objections, because a wrong decision beats no decision.
Between those I run a loop: size-up (impact now, who is here), stabilize (anyone changing things uncoordinated, the best mitigation owned and time-boxed), update (has everyone heard within the cadence), verify (did the last action do what we expected).
Also asked: How do you stop several people changing production at once? · What is the difference between the ops lead and the incident commander? · What do you do when someone more senior joins the incident?
Learn it: 31.4 Roles: the incident commander coordinates, does not debug
How do you communicate during a major incident? Mid
Four audiences, one set of facts, different words: the internal channel (what is going on, who is on it, do you need me), the public status page (is it broken for me, when should I try again), executives (business impact, is it under control, do you need anything from me) and support (what do I tell a customer, is there a workaround).
Every update has the same shape: impact in the reader's words, the action we are taking, and the next update time in UTC. I never promise a fix time; I promise the next update, and post on the cadence even when nothing changed - SEV1 and SEV2 every 30 minutes in our policy. If I lengthen it for a long incident, I say so in an update.
On the status page I move through Investigating, Identified, Monitoring and Resolved, and set each component's status. Outside the channel: no jargon, no guesses at the cause, and no person named as the cause anywhere. A comms lead owns this so the responders are left alone.
Also asked: What would you post on the status page in the first five minutes? · How do you handle an executive who keeps asking for an ETA? · What goes into a good incident update template?
Learn it: 31.8 Communication: cadence, audiences and templates
How do you hand over incident command at the end of a shift? Mid
A handoff is where information gets lost, so it has a fixed form - ICS calls it transfer of command: a briefing with everything needed to continue, then telling everyone who is in charge.
The briefing has five parts:
- impact right now, not at the start;
- what we know: the current hypothesis and what is ruled out;
- everything in flight, each with its owner by name;
- when the next update is due, and to whom;
- open decisions and the risks to watch.
The new IC confirms back in their own words, and I do not leave until they have explicitly accepted: "You're now the incident commander, okay?" Then the role changes in the tool and the channel hears it: "IC is now Casey; next update 16:55 UTC by Casey."
I plan my replacement before I need it. Long incidents rotate the IC every few hours, the way ICS plans in operational periods, because tired people make worse decisions and do not notice it themselves.
Also asked: When would you escalate instead of waiting for the on-call to answer? · What is the difference between functional and hierarchical escalation? · How do you keep on-call sustainable for a team?
Learn it: 31.15 Paging, escalation and handoffs
How do you make decisions during an incident when you do not know the root cause yet? Mid
Mitigate first, understand later: stop the bleeding, restore service, preserve the evidence. I look for the cheapest action that makes users hurt less and that I can undo - generic mitigations like rolling back, turning a flag off, draining a bad region or scaling out.
If a change shortly preceded the impact, the rollback bias says roll it back without proving it caused the problem, after three checks: is it safe (schema migrations), has the path been used, does the timing really match.
I sort options into two-way doors, decided fast, and one-way doors - a restore that loses data, a locking schema change - where I stop and often escalate. I keep several hypotheses, not one theory, and test the cheapest first.
Every significant decision goes on the record with its reason, an owner, a check ("p99 under 800 ms within 10 minutes") and an abort condition, for example incident decide "roll back checkout to 3.0.1" --why "..." in the lab. Then I propose it and poll for strong objections.
Also asked: When is rolling back not the right first move? · How do you avoid fixating on one theory during an incident? · What do you write down when you make a decision in the middle of an incident?
Learn it: 31.21 Deciding under uncertainty: mitigate first, rollback bias, OODA
How do you run a blameless postmortem that actually changes something? Mid
The owner drafts it before the meeting from the timeline, not from memory. Someone who was not the IC facilitates, within a week, about an hour: ground rules, walk the timeline asking at each decision what people knew and expected, contributing factors and what went well, then action items agreed in the room.
Blameless because people who expect punishment leave things out. I look for the second story - what made the action easy, likely or invisible - and watch for hindsight bias and the fundamental attribution error. I ask "what made deploying then look like the right call?", never "why did you deploy on a Friday?", and refer to people by role. If a manager asks who broke it, I answer the need: it will not happen again, and here is the owner of the fix.
No single root cause: contributing factors, plural, until we reach ones we can change. Accountability moves to the action items: each with one owner, a due date, a done-when, a type (prevent, detect, mitigate, process) and a ticket, tracked until done. "Improve monitoring" does not count.
Also asked: What do you say when a manager asks who broke production? · How do you make sure postmortem action items actually get done? · What is the difference between mitigated and resolved?
Learn it: 31.23 Ending the incident, and the postmortem that changes something
How would you measure incident response? Mid
Not by MTTR alone. First I say which R I mean, because repair, recovery, restore and resolve are different clocks. Then:
- Distributions, not means. Durations are skewed: one 290-minute vendor outage turned a median of 17 minutes into a mean of 40 in our quarter. I report the median, the 90th percentile and the longest incidents by name.
- Where the time went, per incident: detect, acknowledge, coordinate, mitigate. Slow detection wants better alerts, slow acks a better rota, slow coordination incident command, slow mitigation better rollbacks.
- User impact: error budget spent, not duration; a 4-minute total outage and a 40-minute SEV3 are not comparable.
- On-call load: pages per shift and how many were actionable, pages out of hours, time on incidents against the 25% line.
- The reviews, because counts are shallow data.
And I never make one of these a target: Goodhart's law - target MTTR and incidents get closed early; target fewer incidents and people stop declaring.
Also asked: Why can MTTR improve when the team changed nothing? · What are the DORA metrics of software delivery? · How would you tell that an on-call rota is burning people out?
Learn it: 31.28 Incident metrics and their traps; on-call health
What does DORA mean for incident response at a bank? Mid
DORA is Regulation (EU) 2022/2554, applied since 17 January 2025. Article 17 requires an incident management process, Article 18 classification, Article 19 reporting major ICT-related incidents to the competent authority.
Classification (RTS 2024/1772): major if it affected critical services and either a successful malicious unauthorised access that may lose data, or two or more materiality thresholds - for example more than 10% of a service's clients, a critical service down over 2 hours, two or more Member States, over EUR 100,000. Recurring incidents with the same apparent root cause can add up to one major.
Clocks (RTS 2025/301): initial notification within 4 hours of classification and no later than 24 hours after awareness; intermediate within 72 hours; final within a month. The weekend relief, noon of the next working day, does not apply to credit institutions, so for a bank a Saturday deadline is a Saturday deadline.
The IC usually does not file. My job is accurate facts and timestamps: awareness and classification times on the timeline, compliance in the room within the first hour, affected clients in the comms plan, a postmortem that feeds the final report.
Also asked: How is a DORA major incident different from your internal severity? · Which timestamps must an incident timeline capture for regulatory reporting? · What other reporting clocks can run during the same incident?
Learn it: 31.31 Regulated environments: DORA major ICT incident reporting
Practise these answers with flashcards and labs Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.