OnCallReady

Azure III: Monitor, KQL & Cost: interview questions

The question you are most likely to get for each topic, a model answer, and what else comes up. From chapter 24 of the course.

How would you investigate a production incident on AKS using Log Analytics? Mid

I run the same six query shapes every time, in UTC, against the workspace (az monitor log-analytics query -w "$W" or the portal):

  1. What is broken? AppRequests per AppRoleName: count, countif(Success == false), percentile(DurationMs, 95), error % with * 100.0.
  2. Since when? The same per bin(TimeGenerated, 1m), where err > 0, top 1 by TimeGenerated asc - the incident start.
  3. What disappeared? Heartbeat | summarize last = max(TimeGenerated) by Computer | where last < ago(5m) - things that stopped reporting never show up as errors.
  4. What else happened then? KubeEvents in a window around the start: NodeNotReady, PreemptScheduled, BackOff. The cause is usually minutes before the symptom.
  5. What changed? AzureActivity writes: who changed what.
  6. Who is affected? dcount(ClientIP) and the pods/endpoints involved.

Joins (pods to KubeNodeInventory, e.g. "were the failing pods on spot nodes?") connect symptom to cause. Remember ingestion latency: the last 1-3 minutes are incomplete.

Also asked: What are the main Log Analytics tables for monitoring an AKS cluster? · How would you reduce the cost of an AKS cluster? · How do you write a good log-based alert?

What is Azure Monitor, and where do the logs for an AKS platform come from? Mid

Azure Monitor is Azure's monitoring umbrella with three stores: Metrics (numeric time series, cheap, 93 days - dashboards and alerts on known signals), Logs (a Log Analytics workspace: tables queried with KQL - where you go during an incident) and the Activity log (control-plane "who did what", 90 days).

Nothing arrives by itself. Sources:

Usually one workspace per environment, timestamps in UTC, billed per GB ingested, and teams get resource-context access to only their cluster's rows.

Also asked: How do you give application teams access to their logs in a shared workspace without exposing everyone's data? · What would you configure on day one of a new AKS platform so that future incidents can be investigated? · Why are the most recent minutes of logs incomplete during an incident?

Learn it: 24.1 Azure Monitor: where the logs actually are

Write a KQL query to find the error logs from one namespace in the last 30 minutes. Mid

ContainerLogV2
| where TimeGenerated > ago(30m)
| where PodNamespace == "shop"
| where LogLevel == "error"
| project TimeGenerated, PodName, LogMessage
| order by TimeGenerated desc

A table, then operators separated by |, like a bash pipeline of rows.

What makes it right:

Also asked: A KQL query returns zero rows but you are sure the data exists. How do you debug it? · What is the difference between has and contains in KQL? · How do you write KQL queries that perform well on a large workspace?

Learn it: 24.3 KQL 1: pipes, filters and strings

How do you calculate the error rate and p95 latency per service in KQL? Mid

AppRequests
| where TimeGenerated > ago(1h)
| summarize requests = count(), errors = countif(Success == false),
            p95 = percentile(DurationMs, 95) by AppRoleName
| extend errorRate = round(errors * 100.0 / requests, 2)
| order by errorRate desc

Also asked: Why can average latency be misleading, and what do you use instead? · How do you get the current state of each pod from a table of per-minute snapshots? · How would you chart errors per minute for one service?

Learn it: 24.5 KQL 2: summarize, bin, percentiles

What are the join kinds in KQL, and what is the trap in the default? Mid

A | join kind=<kind> (B) on Key matches rows of the left side with rows of the right side that have the same key.

Practice: keep the right side small and pre-reduced (summarize arg_max(TimeGenerated, Labels) by Computer in a let), use $left.A == $right.B when key names differ, project away the 1-suffixed duplicates, and compare row counts before and after. union is the other combinator: it stacks tables instead of putting them side by side.

Also asked: How would you find which node pool the failing pods of a service were running on? · What is the difference between join and union in KQL? · How do you keep a KQL join fast on large tables?

Learn it: 24.9 KQL 3: join, union and let for tables

How would you extract a numeric field from log lines in KQL and aggregate it? Mid

Filter first, then cut the field out, typed, then aggregate:

ContainerLogV2
| where TimeGenerated > ago(1h) and LogMessage has "duration_ms"
| extend ms = toint(extract(@"duration_ms=(\d+)", 1, tostring(LogMessage)))
| where isnotnull(ms)
| summarize percentiles(ms, 50, 95, 99) by PodNamespace

Regex over every line of a day is the query that times out - has first. And structured JSON logging beats any clever regex.

Also asked: Why is structured logging valuable for operations, and how does it change your queries? · When would you use parse instead of extract in KQL? · How do you read a value out of a JSON column in KQL?

Learn it: 24.11 KQL 4: getting fields out of text and JSON

What queries would you run in the first ten minutes of an incident? Mid

Six questions, six saved query shapes:

  1. What is broken? Golden signals per service from AppRequests: requests, countif(Success == false), p95, error %.
  2. Since when? Errors per bin(TimeGenerated, 1m), where err > 0 | top 1 by TimeGenerated asc - the start, in UTC, for the timeline.
  3. What disappeared? Heartbeat | summarize last = max(TimeGenerated) by Computer | where last < ago(5m).
  4. What else happened then? KubeEvents from 5 minutes before to 15 after the start: NodeNotReady, PreemptScheduled, TriggeredScaleUp, warnings.
  5. What changed? AzureActivity successful writes: Caller, OperationNameValue, _ResourceId. "Nothing changed" is a claim; this is evidence.
  6. Who is affected? dcount(ClientIP) and make_set(AppRoleInstance) per endpoint - the blast radius.

They live as files in the runbook repo (--analytics-query @kql/golden-signals.kql) or a query pack, so nobody invents queries at 3 a.m.

Also asked: How do you write a good log-based alert? · How do you find what changed in Azure just before an incident? · How do you work out the blast radius of an incident from logs?

Learn it: 24.14 KQL 5: the incident playbook

How would you reduce the cost of an AKS cluster? Mid

First know where it goes: compute (node VMs, 60-80%), logs (Log Analytics ingestion - the surprise), storage, network, the cluster tier.

Compute:

Logs: find the biggest tables with Usage | summarize GB = sum(Quantity) / 1024 by DataType, then fix the log level at the source, filter at collection (DCR), Basic plan for noisy rarely-queried tables, a daily cap on dev only.

Orphans: az disk list --query "[?diskState=='Unattached']", public IPs with no ipConfiguration, old snapshots.

Then keep it fixed: budgets with forecast alerts, mandatory tags for showback per team.

Also asked: What are the main cost components of running AKS? · How do you find orphaned Azure resources that are still being billed? · How do you make teams aware of what their workloads cost?

Learn it: 24.26 Cost: where the money goes in an AKS platform

Practise these answers with flashcards and labs Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.