Azure III: Monitor, KQL & Cost: interview questions
The question you are most likely to get for each topic, a model answer, and what else comes up. From chapter 24 of the course.
How would you investigate a production incident on AKS using Log Analytics? Mid
I run the same six query shapes every time, in UTC, against the workspace (az monitor log-analytics query -w "$W" or the portal):
- What is broken?
AppRequestsperAppRoleName: count,countif(Success == false),percentile(DurationMs, 95), error % with* 100.0. - Since when? The same per
bin(TimeGenerated, 1m),where err > 0,top 1 by TimeGenerated asc- the incident start. - What disappeared?
Heartbeat | summarize last = max(TimeGenerated) by Computer | where last < ago(5m)- things that stopped reporting never show up as errors. - What else happened then?
KubeEventsin a window around the start:NodeNotReady,PreemptScheduled,BackOff. The cause is usually minutes before the symptom. - What changed?
AzureActivitywrites: who changed what. - Who is affected?
dcount(ClientIP)and the pods/endpoints involved.
Joins (pods to KubeNodeInventory, e.g. "were the failing pods on spot nodes?") connect symptom to cause. Remember ingestion latency: the last 1-3 minutes are incomplete.
Also asked: What are the main Log Analytics tables for monitoring an AKS cluster? · How would you reduce the cost of an AKS cluster? · How do you write a good log-based alert?
What is Azure Monitor, and where do the logs for an AKS platform come from? Mid
Azure Monitor is Azure's monitoring umbrella with three stores: Metrics (numeric time series, cheap, 93 days - dashboards and alerts on known signals), Logs (a Log Analytics workspace: tables queried with KQL - where you go during an incident) and the Activity log (control-plane "who did what", 90 days).
Nothing arrives by itself. Sources:
- Container Insights (the AKS add-on, an agent on every node):
ContainerLogV2(container stdout/stderr),KubePodInventoryandKubeNodeInventory(snapshots every minute),KubeEvents,Perf,Heartbeat. - Application Insights:
AppRequests,AppDependencies,AppExceptionsfrom an SDK in the app. - Diagnostic settings per resource: Key Vault audit, App Gateway access logs, AKS control-plane logs - into
AzureDiagnosticsor resource-specific tables. Off by default: a vault without one keeps no record of who read which secret. - Activity log export:
AzureActivity.
Usually one workspace per environment, timestamps in UTC, billed per GB ingested, and teams get resource-context access to only their cluster's rows.
Also asked: How do you give application teams access to their logs in a shared workspace without exposing everyone's data? · What would you configure on day one of a new AKS platform so that future incidents can be investigated? · Why are the most recent minutes of logs incomplete during an incident?
Write a KQL query to find the error logs from one namespace in the last 30 minutes. Mid
ContainerLogV2
| where TimeGenerated > ago(30m)
| where PodNamespace == "shop"
| where LogLevel == "error"
| project TimeGenerated, PodName, LogMessage
| order by TimeGenerated desc
A table, then operators separated by |, like a bash pipeline of rows.
What makes it right:
- Time first - the workspace stores data in time chunks; a time filter skips most of it. Without one the query scans the whole retention.
- Cheap, selective filters next (
==on a column),projectearly to keep rows narrow. - String operators:
==is case-sensitive,=~is not.hasmatches a whole term using the index - the default choice;containsscans for any substring.has "status=5"finds nothing, because the terms arestatusand503. - Types:
ResultCodeinAppRequestsis a string -startswith "5"ortoint(), not>= 500. takereturns arbitrary rows; "latest N" istop N by TimeGenerated.
Also asked: A KQL query returns zero rows but you are sure the data exists. How do you debug it? · What is the difference between has and contains in KQL? · How do you write KQL queries that perform well on a large workspace?
Learn it: 24.3 KQL 1: pipes, filters and strings
How do you calculate the error rate and p95 latency per service in KQL? Mid
AppRequests
| where TimeGenerated > ago(1h)
| summarize requests = count(), errors = countif(Success == false),
p95 = percentile(DurationMs, 95) by AppRoleName
| extend errorRate = round(errors * 100.0 / requests, 2)
| order by errorRate desc
summarize ... bygroups the rows - one row out per service - and computes each aggregate per group (sort | uniq -cgrown up). Name the aggregates; they become your columns and alert fields.- The integer-division trap:
count()returns a long, anderrors / requestsis 0. One real operand (100.0) fixes it - otherwise an alert on it never fires. - Percentiles, not averages: the average hides the slow tail users feel. KQL percentiles are estimates - fine for operations.
- For a trend, add
bin(TimeGenerated, 5m)to thebyandrender timechart; empty buckets do not appear, so a gap is "no data", not zero. - For the current state from snapshot tables:
summarize arg_max(TimeGenerated, PodStatus) by Namespace, Name.
Also asked: Why can average latency be misleading, and what do you use instead? · How do you get the current state of each pod from a table of per-minute snapshots? · How would you chart errors per minute for one service?
Learn it: 24.5 KQL 2: summarize, bin, percentiles
What are the join kinds in KQL, and what is the trap in the default? Mid
A | join kind=<kind> (B) on Key matches rows of the left side with rows of the right side that have the same key.
inner- every matching pair.leftouter- every left row; right columns empty where nothing matched. The safe choice for enrichment - nothing silently disappears.leftanti- left rows with no match ("pods with no node row").leftsemi- left rows that have a match, left columns only.innerunique- the default: deduplicates the left side on the key, then inner-joins. Join 60 per-minute pod snapshots with the default and you get one row per pod - so always writekind=.
Practice: keep the right side small and pre-reduced (summarize arg_max(TimeGenerated, Labels) by Computer in a let), use $left.A == $right.B when key names differ, project away the 1-suffixed duplicates, and compare row counts before and after. union is the other combinator: it stacks tables instead of putting them side by side.
Also asked: How would you find which node pool the failing pods of a service were running on? · What is the difference between join and union in KQL? · How do you keep a KQL join fast on large tables?
How would you extract a numeric field from log lines in KQL and aggregate it? Mid
Filter first, then cut the field out, typed, then aggregate:
ContainerLogV2
| where TimeGenerated > ago(1h) and LogMessage has "duration_ms"
| extend ms = toint(extract(@"duration_ms=(\d+)", 1, tostring(LogMessage)))
| where isnotnull(ms)
| summarize percentiles(ms, 50, 95, 99) by PodNamespace
extract(regex, group, text)- one field by regex, returns a string; convert withtoint(). Verbatim strings (@"...") keep the backslashes.parse LogMessage with * "status=" status:int " duration_ms=" ms:long ...- several fields at once when every line has the same layout; typed captures let you compare numerically. Lines that do not match give empty captures - filter them out.- JSON:
LogMessageis dynamic, so JSON logs can be walked directly (LogMessage.level); JSON in string columns needsparse_json(Labels)[0].agentpoolandtostring().
Regex over every line of a day is the query that times out - has first. And structured JSON logging beats any clever regex.
Also asked: Why is structured logging valuable for operations, and how does it change your queries? · When would you use parse instead of extract in KQL? · How do you read a value out of a JSON column in KQL?
What queries would you run in the first ten minutes of an incident? Mid
Six questions, six saved query shapes:
- What is broken? Golden signals per service from
AppRequests: requests,countif(Success == false), p95, error %. - Since when? Errors per
bin(TimeGenerated, 1m),where err > 0 | top 1 by TimeGenerated asc- the start, in UTC, for the timeline. - What disappeared?
Heartbeat | summarize last = max(TimeGenerated) by Computer | where last < ago(5m). - What else happened then?
KubeEventsfrom 5 minutes before to 15 after the start:NodeNotReady,PreemptScheduled,TriggeredScaleUp, warnings. - What changed?
AzureActivitysuccessful writes:Caller,OperationNameValue,_ResourceId. "Nothing changed" is a claim; this is evidence. - Who is affected?
dcount(ClientIP)andmake_set(AppRoleInstance)per endpoint - the blast radius.
They live as files in the runbook repo (--analytics-query @kql/golden-signals.kql) or a query pack, so nobody invents queries at 3 a.m.
Also asked: How do you write a good log-based alert? · How do you find what changed in Azure just before an incident? · How do you work out the blast radius of an incident from logs?
Learn it: 24.14 KQL 5: the incident playbook
How would you reduce the cost of an AKS cluster? Mid
First know where it goes: compute (node VMs, 60-80%), logs (Log Analytics ingestion - the surprise), storage, network, the cluster tier.
Compute:
- Right-size requests - the scheduler and autoscaler work from requests, not usage; requests at twice real usage mean twice the nodes. Compare with
Perf/InsightsMetrics. - Make sure the cluster autoscaler can scale down - bare pods, local storage and too-strict PDBs pin nodes.
- Spot pools for interruptible work,
az aks stopfor dev clusters at night. - Commitments (reserved instances, savings plans) for the steady floor, never the peak.
Logs: find the biggest tables with Usage | summarize GB = sum(Quantity) / 1024 by DataType, then fix the log level at the source, filter at collection (DCR), Basic plan for noisy rarely-queried tables, a daily cap on dev only.
Orphans: az disk list --query "[?diskState=='Unattached']", public IPs with no ipConfiguration, old snapshots.
Then keep it fixed: budgets with forecast alerts, mandatory tags for showback per team.
Also asked: What are the main cost components of running AKS? · How do you find orphaned Azure resources that are still being billed? · How do you make teams aware of what their workloads cost?
Learn it: 24.26 Cost: where the money goes in an AKS platform
Practise these answers with flashcards and labs Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.