At 3 a.m. with checkout down you do not want to invent queries. You want to run the six you already have. Most incident questions - what, since when, what vanished, what else happened, what changed, who is hurt - are the same six query shapes every time. Learn them as patterns, keep them in files, and adapt the table and column names.
What you need to know already: 24.3-24.11 (filters, summarize, bin, join, let, parse), 0.1 (golden signals), 0.21 (alerting philosophy), 0.29 (running an incident, the timeline), 22.13 (RBAC and the control plane).
1. What is broken? (golden signals by service)
AppRequests
| where TimeGenerated > ago(1h)
| summarize req = count(), err = countif(Success == false), p95 = percentile(DurationMs, 95) by AppRoleName
| extend errPct = round(err * 100.0 / req, 2)
| order by errPct desc
2. Since when? (a time series, then the first bad bucket)
AppRequests
| where TimeGenerated > ago(3h) and AppRoleName == "payments-api"
| summarize err = countif(ResultCode startswith "5"), p95 = percentile(DurationMs, 95) by bin(TimeGenerated, 1m)
| where err > 0
| top 1 by TimeGenerated asc
Per minute, count the 5xx; keep only minutes with errors; the earliest one (top 1 by TimeGenerated asc) is when it started. Coarse bins to see the shape (5m), fine bins to pin the start (1m). The first bad minute is your incident start in the timeline - in UTC.
3. What disappeared? (last seen)
Heartbeat
| summarize last = max(TimeGenerated) by Computer
| where last < ago(5m)
For each node, the newest heartbeat; keep the nodes whose newest one is older than 5 minutes - they went silent. Things that stop reporting do not show up in "what is erroring" queries. Ask for the last time each thing was seen.
4. What else happened then? (correlate by time window)
let start = datetime(2026-09-24 09:20);
KubeEvents
| where TimeGenerated between ((start - 5m) .. (start + 15m))
| where KubeEventType == "Warning" or Reason in ("NodeNotReady", "PreemptScheduled", "TriggeredScaleUp")
| project TimeGenerated, ObjectKind, Name, Reason, Message
| order by TimeGenerated asc
KubeEventType == "Warning" keeps Kubernetes' warning events (Ch 15). The three reasons named are ones you will meet in the incident: NodeNotReady (a node stopped answering), PreemptScheduled (Azure announced a spot eviction, 23.25), TriggeredScaleUp (the cluster autoscaler is adding a node, 23.25).
A window around the start, across the tables that describe the platform (events, node inventory, pod inventory), read in time order. The cause is usually a few minutes before the symptom.
5. What changed? (the control plane)
AzureActivity
| where TimeGenerated > ago(3d)
| where ActivityStatusValue == "Success" and OperationNameValue has "WRITE"
| project TimeGenerated, Caller, OperationNameValue, _ResourceId
| order by TimeGenerated desc
AzureActivity is the Activity log (24.1): every change made through Azure's control plane, with Caller (who), OperationNameValue (what, e.g. MICROSOFT.AUTHORIZATION/ROLEASSIGNMENTS/WRITE) and _ResourceId (on which resource). Deployments, role assignments, scale operations, config changes. "Nothing changed" is a claim; this is the evidence.
6. Who is affected? (blast radius)
The blast radius is how much of the system and how many users a failure hit. Here: per service and endpoint, the failed requests, how many distinct client IPs saw them, and which pods served them.
AppRequests
| where TimeGenerated between (start .. (start + 15m)) and Success == false
| summarize failed = count(), users = dcount(ClientIP), pods = make_set(AppRoleInstance) by AppRoleName, Name
Alert queries
You met alerting as an idea in 0.21. In Azure, a log search alert (an alert rule of type "log search") is a KQL query plus a threshold that Azure runs on a schedule, say every 5 minutes; when the threshold is crossed it notifies people through an action group (a saved list of who to email, text or call). The query:
AppRequests
| where TimeGenerated > ago(5m) and AppRoleName == "payments-api"
| summarize errPct = countif(Success == false) * 100.0 / count()
| where errPct > 5
Make the query return rows only when something is wrong, alert on "number of results > 0", and test it against a past incident window before trusting it. Note the 100.0: the integer-division bug turns alerts into ones that never fire.
Save them
Keep these in a kql/ folder in the team's runbook repo (a runbook is the written "what to do when X breaks" guide), one file per question, and run them with --analytics-query @kql/golden-signals.kql. In the portal, save them to a query pack (an Azure resource that holds shared saved queries) so everyone on call has the same ones.
What you can now do
- Answer the six incident questions with six query shapes: what, since when, what vanished, what else happened, what changed, who is affected.
- Write an alert query that returns rows only when something is wrong.
- Keep a team's queries in files and run them with
@file.