Checkout is failing. Six services run on 40 pods spread over five nodes of an AKS cluster. You could kubectl logs each pod one by one (15.42) - and the pod that crashed ten minutes ago is gone, its logs with it. What you want is every log line from every pod, kept in one place, searchable after the fact. On Azure that place is a Log Analytics workspace, and this chapter teaches you to search it.
What you need to know already: 22.1-22.7 (az, -o table, --query, variables from -o tsv), 22.13 (RBAC roles and scopes), 23.20 (AKS from the Azure side), 15.14 (pods), 15.42 (kubectl logs), 0.1 (the golden signals), 0.21 (alerting).
Three words first
- Log - one line of text a program writes about something that happened:
GET /api/cart 200 41ms,connection refused. You met them in journald (2.14) and nginx's access log (Ch 7). - Metric - a number measured again and again over time: CPU percent every minute, requests per second. A list of (time, value) pairs is a time series.
- Monitoring - collecting logs and metrics from everything you run, so you can see what it is doing and find out why it broke.
Azure Monitor is Azure's name for all of its monitoring features together. It keeps data in three separate stores:
Metrics numeric time series, 1-minute grain, 93 days, cheap, fast "CPU of this VM"
Logs a Log Analytics workspace: tables of records, queried with KQL "every 5xx and its pod"
Activity log who did what to which resource (control plane), 90 days "who deleted the disk"
Reading the table: grain is how often a value is recorded; 93 days / 90 days is how long Azure keeps it; the control plane (22.23) is the API that creates, changes and deletes resources, as opposed to the data inside them.
Metrics are for dashboards (screens of charts) and for alerts (a rule that notifies a person when a number crosses a line) on signals you already know to watch. Logs are where you go during an incident, because you can ask questions nobody thought of in advance. This chapter is about logs.
The workspace
A Log Analytics workspace is an Azure resource (it lives in a resource group like everything else, 22.1) that stores log records in tables - rows and named columns, like a spreadsheet or a SQL table. You read it with KQL (Kusto Query Language), Azure's query language, which the next lessons teach. Ours is called law-sysop (LAW = Log Analytics workspace).
az monitor log-analytics workspace show -g rg-oncall-lab -n law-sysop --query ...: show one workspace; -g its resource group; -n its name; --query a JMESPath expression (22.4) that picks four fields and renames them.
$ az monitor log-analytics workspace show -g rg-oncall-lab -n law-sysop --query "{name:name, id:customerId, retention:retentionInDays, sku:sku.name}"
{
"id": "b5f3c1d2-7a4e-4f7b-9c1e-2d6a8f0e4b31",
"name": "law-sysop",
"retention": 30,
"sku": "PerGB2018"
}
What each field means:
id(calledcustomerIdin the full output) - a GUID (a random-looking unique ID in the form 8-4-4-4-12 hex digits) that the query API uses to find the workspace. It is not the ARM resource ID (/subscriptions/.../workspaces/law-sysop, 22.1) and not the name - you will need exactly this one.retention- days a record is kept before it is deleted.sku- the pricing plan (SKU = stock-keeping unit, Azure's word for a pricing tier).PerGB2018= pay per gigabyte you send in.
Most teams run one workspace per environment (or per region) that every resource sends to, so one query can combine a pod's logs with the Key Vault's audit log with the load balancer's records.
What writes into it
Data does not appear by itself. Something has to send it:
- Container Insights - the AKS monitoring add-on (an optional feature you switch on per cluster). An agent (a small program, here the Azure Monitor Agent, running on every node) collects and uploads:
ContainerLogV2(stdout/stderr of every container, Ch 10),KubePodInventory(a snapshot of every pod every minute),KubeNodeInventory,KubeEvents(the eventskubectl get eventsshows, Ch 15),Perf/InsightsMetrics(CPU, memory),Heartbeat(an "I am alive" row per node per minute). - Application Insights - records written by the application code itself, one row per request it served (
AppRequests), per call it made to something else (AppDependencies), per exception (AppExceptions) and per log call (AppTraces). The app sends them through a library added to its code (an SDK, software development kit). - Diagnostic settings - a setting on an Azure resource that says "send your own logs to this workspace": Key Vault audit (who read which secret), App Gateway access logs (23.16), NSG flow logs (23.3), AKS control plane logs (
kube-audit,kube-apiserver...). They land either in the old sharedAzureDiagnosticstable - where column names get a type suffix:_sstring,_dnumber,_gGUID,_btrue/false, as inhttpStatusCode_d,identity_claim_oid_g- or, for services that support it, in their own resource-specific tables (AKSAudit,AGWAccessLogs...). Prefer resource-specific for new setups: proper column names, one table per service. - Activity log export:
AzureActivity(the control-plane "who did what").
Nothing is logged by default except the Activity log. A Key Vault without a diagnostic setting keeps no record of who read which secret - set it up before the incident, not after.
az monitor diagnostic-settings create: create a diagnostic setting; -n its name; --resource the ID of the resource whose logs to send (here pasted in by $(az keyvault show ... --query id -o tsv), 22.7); --workspace where to send them; --logs a JSON list of log categories to switch on (AuditEvent = Key Vault's "who did what" log).
az monitor diagnostic-settings create -n kv-audit --resource $(az keyvault show -n kv-shop-prod --query id -o tsv) \
--workspace $(az monitor log-analytics workspace show -g rg-oncall-lab -n law-sysop --query id -o tsv) \
--logs '[{"category":"AuditEvent","enabled":true}]'
Querying from the CLI
Two commands, run once per shell and then for every query:
W=$(az monitor log-analytics workspace show ... --query customerId -o tsv): put the workspace GUID in a shell variableW(-o tsvprints the bare value with no quotes, 22.3).az monitor log-analytics query -w "$W" --analytics-query '<KQL>' -o table: run one KQL query;-wwhich workspace (the GUID);--analytics-querythe query text, in single quotes so bash leaves the|and"inside alone (Ch 6);-o tableprint it as a table.
The query below reads: take the Heartbeat table, and for each Computer (node) keep the latest TimeGenerated (the time the row was written) as last. You will learn summarize in 24.5; for now read it as "group by".
$ W=$(az monitor log-analytics workspace show -g rg-oncall-lab -n law-sysop --query customerId -o tsv)
$ az monitor log-analytics query -w "$W" --analytics-query 'Heartbeat | summarize last=max(TimeGenerated) by Computer' -o table
Computer TableName Last
------------------------------ ------------- --------------------
aks-system-31415926-vmss000000 PrimaryResult 2026-09-24T09:41:07Z
...
Things to know about the CLI output:
- the extra
TableNamecolumn (PrimaryResult) is added by the CLI; ignore it -o tablesorts the columns by name (capitals first) and capitalises the first letter: yourlastshows up asLast, afterTableName. With a--query(--query "[].{Computer:Computer, last:last}") it keeps your order- every value comes back as a string (
"count_": "12") in JSON output; use--query/jq(22.4, Ch 7) with that in mind -wwants the customerId; the workspace name fails withPathNotFoundError--timespan PT1Hrestricts every table to the last hour, on top of any time filter in the query.PT1His the ISO 8601 way to write a duration: P = period, T = the time part follows, 1H = one hour (PT30M,P1D)- long queries are unreadable inline: put them in a file and pass
--analytics-query @hunt.kql- the@filesyntax works for anyazargument
The portal (Azure's website, 22.17) has a Logs blade (a blade is a page of the portal) that runs the same queries with a chart button; the CLI is what you script, alert-test and paste into incident channels.
Timestamps are UTC
TimeGenerated is UTC (Coordinated Universal Time, the same clock everywhere in the world, no summer time), always. The portal can display local time; the data is UTC. Write incident timelines in UTC and you will never be an hour off during summer time again.
Cost, briefly (the cost lesson has the rest)
Workspaces bill by GB ingested - ingestion is data arriving in the workspace, so you pay for every gigabyte sent in (and for retention past the free period). The Usage table records GB per table per day, and every record carries _BilledSize. A single deployment left at debug log level can multiply the bill - there is a mission about exactly that.
A tour of the Container Insights tables
KubePodInventory one row per pod per minute: Namespace, Name, PodStatus, Computer (node),
PodRestartCount (cumulative!), ContainerLastStatus, ControllerName
KubeNodeInventory one row per node per minute: Status (Ready/NotReady), Labels (JSON), KubeletVersion
KubeEvents Kubernetes events: Reason (BackOff, FailedMount, FailedScheduling...), Message, ObjectKind
ContainerLogV2 every stdout/stderr line: PodName, PodNamespace, ContainerName, LogMessage (dynamic), LogLevel
Perf CPU/memory per node and container: ObjectName, CounterName, CounterValue
Heartbeat one row per agent per minute - "is this node still reporting?"
Some words in that table: PodRestartCount is how many times the pod's container has been restarted (Ch 15), and cumulative means it only ever grows - it is a running total, not "restarts this minute". dynamic is KQL's type for a JSON value (lesson 24.11).
Inventory tables are snapshots, not event streams: a pod that ran all hour has sixty rows, one per minute. That shapes every query you write against them: you will want the latest row per thing, the growth of a counter (max - min), first and last seen (the smallest and largest TimeGenerated). The next lessons teach the operators for each.
Ingestion latency
Ingestion latency is the delay between something happening and its record being queryable. Records typically arrive 1-3 minutes after the event (more under load). During an incident the last couple of minutes are incomplete; an alert (0.21) that looks back only 1 minute over a table that is 3 minutes late misses things. The column _TimeReceived (real workspaces; not simulated here) shows the delay.
Access
Querying is an RBAC action like any other (22.13): it needs Microsoft.OperationalInsights/workspaces/query/*, which the built-in role Log Analytics Reader on the workspace grants. With resource-context access, a team with Reader on its own AKS cluster can query only the rows that came from that cluster, without rights on the whole workspace - the usual way to share one workspace safely.
What you can now do
- Say what Azure Monitor's three stores hold and which one you open during an incident.
- Name the tables Container Insights fills and what one row of each means.
- Query a workspace from the terminal with
az monitor log-analytics query -w "$W".