OnCallReady

Lesson 24.14 · Azure III: Monitor, KQL & Cost · 13 min read

KQL 5: the incident playbook

In plain words

Imagine a fire brigade with a laminated checklist in every truck: Where is the fire? When did it start? Is anyone missing? What was going on just before? Did anyone change something in the building? Who needs help? They don't invent the questions at each fire; they run the same list every time and fill in the answers.

The incident playbook is that checklist, as six KQL queries: golden signals by service, the first bad minute, what stopped reporting, what else happened in the window, what changed in AzureActivity, and who's affected. You keep them in files in kql/ in the runbook repo and run them with --analytics-query @kql/golden-signals.kql, adapting the table and column names.

At 3 a.m. with checkout down you do not want to invent queries. You want to run the six you already have. Most incident questions - what, since when, what vanished, what else happened, what changed, who is hurt - are the same six query shapes every time. Learn them as patterns, keep them in files, and adapt the table and column names.

What you need to know already: 24.3-24.11 (filters, summarize, bin, join, let, parse), 0.1 (golden signals), 0.21 (alerting philosophy), 0.29 (running an incident, the timeline), 22.13 (RBAC and the control plane).

1. What is broken? (golden signals by service)

AppRequests
| where TimeGenerated > ago(1h)
| summarize req = count(), err = countif(Success == false), p95 = percentile(DurationMs, 95) by AppRoleName
| extend errPct = round(err * 100.0 / req, 2)
| order by errPct desc

2. Since when? (a time series, then the first bad bucket)

AppRequests
| where TimeGenerated > ago(3h) and AppRoleName == "payments-api"
| summarize err = countif(ResultCode startswith "5"), p95 = percentile(DurationMs, 95) by bin(TimeGenerated, 1m)
| where err > 0
| top 1 by TimeGenerated asc

Per minute, count the 5xx; keep only minutes with errors; the earliest one (top 1 by TimeGenerated asc) is when it started. Coarse bins to see the shape (5m), fine bins to pin the start (1m). The first bad minute is your incident start in the timeline - in UTC.

3. What disappeared? (last seen)

Heartbeat
| summarize last = max(TimeGenerated) by Computer
| where last < ago(5m)

For each node, the newest heartbeat; keep the nodes whose newest one is older than 5 minutes - they went silent. Things that stop reporting do not show up in "what is erroring" queries. Ask for the last time each thing was seen.

4. What else happened then? (correlate by time window)

let start = datetime(2026-09-24 09:20);
KubeEvents
| where TimeGenerated between ((start - 5m) .. (start + 15m))
| where KubeEventType == "Warning" or Reason in ("NodeNotReady", "PreemptScheduled", "TriggeredScaleUp")
| project TimeGenerated, ObjectKind, Name, Reason, Message
| order by TimeGenerated asc

KubeEventType == "Warning" keeps Kubernetes' warning events (Ch 15). The three reasons named are ones you will meet in the incident: NodeNotReady (a node stopped answering), PreemptScheduled (Azure announced a spot eviction, 23.25), TriggeredScaleUp (the cluster autoscaler is adding a node, 23.25).

A window around the start, across the tables that describe the platform (events, node inventory, pod inventory), read in time order. The cause is usually a few minutes before the symptom.

5. What changed? (the control plane)

AzureActivity
| where TimeGenerated > ago(3d)
| where ActivityStatusValue == "Success" and OperationNameValue has "WRITE"
| project TimeGenerated, Caller, OperationNameValue, _ResourceId
| order by TimeGenerated desc

AzureActivity is the Activity log (24.1): every change made through Azure's control plane, with Caller (who), OperationNameValue (what, e.g. MICROSOFT.AUTHORIZATION/ROLEASSIGNMENTS/WRITE) and _ResourceId (on which resource). Deployments, role assignments, scale operations, config changes. "Nothing changed" is a claim; this is the evidence.

6. Who is affected? (blast radius)

The blast radius is how much of the system and how many users a failure hit. Here: per service and endpoint, the failed requests, how many distinct client IPs saw them, and which pods served them.

AppRequests
| where TimeGenerated between (start .. (start + 15m)) and Success == false
| summarize failed = count(), users = dcount(ClientIP), pods = make_set(AppRoleInstance) by AppRoleName, Name

Alert queries

You met alerting as an idea in 0.21. In Azure, a log search alert (an alert rule of type "log search") is a KQL query plus a threshold that Azure runs on a schedule, say every 5 minutes; when the threshold is crossed it notifies people through an action group (a saved list of who to email, text or call). The query:

AppRequests
| where TimeGenerated > ago(5m) and AppRoleName == "payments-api"
| summarize errPct = countif(Success == false) * 100.0 / count()
| where errPct > 5

Make the query return rows only when something is wrong, alert on "number of results > 0", and test it against a past incident window before trusting it. Note the 100.0: the integer-division bug turns alerts into ones that never fire.

Save them

Keep these in a kql/ folder in the team's runbook repo (a runbook is the written "what to do when X breaks" guide), one file per question, and run them with --analytics-query @kql/golden-signals.kql. In the portal, save them to a query pack (an Azure resource that holds shared saved queries) so everyone on call has the same ones.

What you can now do

Why it helps

In the first minutes of an incident, the team that already has its questions written down wins. Running the same six shapes every time means you don't forget "what changed?" or "what stopped reporting?", the two questions most often skipped, and your timeline is in UTC from the start. Saved as files and query packs, they're shared by everyone on call, including someone new at 3am.

Alert queries come from the same patterns, with one discipline: return rows only when something is wrong, avoid the integer-division bug, and test against a past incident window. Knowing that is the difference between alerts that catch the next outage and alerts that are either silent or so noisy that everyone ignores them.

FAQ

Why do I need a "what disappeared" query?

Because things that stop reporting produce no rows, so they never show up in queries for errors or high latency. A node that died, an agent that stopped, a service that crashed and isn't being called don't generate error records; they just go quiet. Asking for the last time each thing was seen, like Heartbeat | summarize last = max(TimeGenerated) by Computer | where last < ago(5m), catches exactly what the other queries miss.

How do I look for the cause rather than the symptom?

Look at a window around the incident start, beginning a few minutes before it, across the tables that describe the platform: KubeEvents, node and pod inventory, AzureActivity, deployment history. Read them in time order. Causes usually precede symptoms by minutes: a node preemption, a deployment, a config change, a scale-down. Symptoms are in application tables; causes are mostly in platform tables.

How should a log search alert query be written?

So that it returns rows only when something is wrong, and the alert fires on "number of results greater than 0". Filter a short window, compute the signal with real division, like countif(Success == false) * 100.0 / count(), and end with the threshold in a where. Add a minimum volume so one failure out of two requests doesn't page anyone. Then run it against a past incident window to prove it would have fired.

Why keep queries in a repository?

So they're reviewed, versioned and shared, and on-call doesn't write them from scratch under pressure. One file per question in something like kql/, run with --analytics-query @kql/file.kql, works from any terminal and can be pasted into incident channels. In the portal, a query pack gives everyone the same saved queries. After each incident, improve the files with what you wished you'd had.

Why write incident timelines in UTC?

Because TimeGenerated and all Azure log timestamps are UTC, while people and the portal often show local time, which shifts with summer time. Mixing them puts events an hour or two apart in the timeline and can make a cause appear after its effect. Writing everything in UTC, and converting only for communication outside the team, removes that whole class of mistakes.

In an interview Mid

What queries would you run in the first ten minutes of an incident?

Six questions, six saved query shapes:

  1. What is broken? Golden signals per service from AppRequests: requests, countif(Success == false), p95, error %.
  2. Since when? Errors per bin(TimeGenerated, 1m), where err > 0 | top 1 by TimeGenerated asc - the start, in UTC, for the timeline.
  3. What disappeared? Heartbeat | summarize last = max(TimeGenerated) by Computer | where last < ago(5m).
  4. What else happened then? KubeEvents from 5 minutes before to 15 after the start: NodeNotReady, PreemptScheduled, TriggeredScaleUp, warnings.
  5. What changed? AzureActivity successful writes: Caller, OperationNameValue, _ResourceId. "Nothing changed" is a claim; this is evidence.
  6. Who is affected? dcount(ClientIP) and make_set(AppRoleInstance) per endpoint - the blast radius.

They live as files in the runbook repo (--analytics-query @kql/golden-signals.kql) or a query pack, so nobody invents queries at 3 a.m.

Also asked: How do you write a good log-based alert? · How do you find what changed in Azure just before an incident? · How do you work out the blast radius of an incident from logs?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.