OnCallReady

Lesson 24.1 · Azure III: Monitor, KQL & Cost · 17 min read

Azure Monitor: where the logs actually are

In plain words

Imagine a school that keeps three kinds of records. A thermometer on the wall writes the temperature every minute on a chart: simple numbers, cheap to keep. A big diary in the office where every event is written in detail: who was late, which class had a fire drill, what the nurse saw. And a visitor book at the gate: who came in to change something, and when.

Azure Monitor has the same three: metrics, numbers every minute, kept 93 days; logs in a Log Analytics workspace like law-sysop, detailed records you query with KQL; and the activity log, who changed which resource. The detail that surprises people: apart from the activity log, nothing is written unless you set it up. A Key Vault without a diagnostic setting keeps no record of who read a secret.

Checkout is failing. Six services run on 40 pods spread over five nodes of an AKS cluster. You could kubectl logs each pod one by one (15.42) - and the pod that crashed ten minutes ago is gone, its logs with it. What you want is every log line from every pod, kept in one place, searchable after the fact. On Azure that place is a Log Analytics workspace, and this chapter teaches you to search it.

What you need to know already: 22.1-22.7 (az, -o table, --query, variables from -o tsv), 22.13 (RBAC roles and scopes), 23.20 (AKS from the Azure side), 15.14 (pods), 15.42 (kubectl logs), 0.1 (the golden signals), 0.21 (alerting).

Three words first

Azure Monitor is Azure's name for all of its monitoring features together. It keeps data in three separate stores:

Metrics         numeric time series, 1-minute grain, 93 days, cheap, fast      "CPU of this VM"
Logs            a Log Analytics workspace: tables of records, queried with KQL  "every 5xx and its pod"
Activity log    who did what to which resource (control plane), 90 days          "who deleted the disk"

Reading the table: grain is how often a value is recorded; 93 days / 90 days is how long Azure keeps it; the control plane (22.23) is the API that creates, changes and deletes resources, as opposed to the data inside them.

Metrics are for dashboards (screens of charts) and for alerts (a rule that notifies a person when a number crosses a line) on signals you already know to watch. Logs are where you go during an incident, because you can ask questions nobody thought of in advance. This chapter is about logs.

The workspace

A Log Analytics workspace is an Azure resource (it lives in a resource group like everything else, 22.1) that stores log records in tables - rows and named columns, like a spreadsheet or a SQL table. You read it with KQL (Kusto Query Language), Azure's query language, which the next lessons teach. Ours is called law-sysop (LAW = Log Analytics workspace).

az monitor log-analytics workspace show -g rg-oncall-lab -n law-sysop --query ...: show one workspace; -g its resource group; -n its name; --query a JMESPath expression (22.4) that picks four fields and renames them.

$ az monitor log-analytics workspace show -g rg-oncall-lab -n law-sysop --query "{name:name, id:customerId, retention:retentionInDays, sku:sku.name}"
{
  "id": "b5f3c1d2-7a4e-4f7b-9c1e-2d6a8f0e4b31",
  "name": "law-sysop",
  "retention": 30,
  "sku": "PerGB2018"
}

What each field means:

Most teams run one workspace per environment (or per region) that every resource sends to, so one query can combine a pod's logs with the Key Vault's audit log with the load balancer's records.

What writes into it

Data does not appear by itself. Something has to send it:

Nothing is logged by default except the Activity log. A Key Vault without a diagnostic setting keeps no record of who read which secret - set it up before the incident, not after.

az monitor diagnostic-settings create: create a diagnostic setting; -n its name; --resource the ID of the resource whose logs to send (here pasted in by $(az keyvault show ... --query id -o tsv), 22.7); --workspace where to send them; --logs a JSON list of log categories to switch on (AuditEvent = Key Vault's "who did what" log).

az monitor diagnostic-settings create -n kv-audit --resource $(az keyvault show -n kv-shop-prod --query id -o tsv) \
  --workspace $(az monitor log-analytics workspace show -g rg-oncall-lab -n law-sysop --query id -o tsv) \
  --logs '[{"category":"AuditEvent","enabled":true}]'

Querying from the CLI

Two commands, run once per shell and then for every query:

The query below reads: take the Heartbeat table, and for each Computer (node) keep the latest TimeGenerated (the time the row was written) as last. You will learn summarize in 24.5; for now read it as "group by".

$ W=$(az monitor log-analytics workspace show -g rg-oncall-lab -n law-sysop --query customerId -o tsv)
$ az monitor log-analytics query -w "$W" --analytics-query 'Heartbeat | summarize last=max(TimeGenerated) by Computer' -o table
Computer                        TableName      Last
------------------------------  -------------  --------------------
aks-system-31415926-vmss000000  PrimaryResult  2026-09-24T09:41:07Z
...

Things to know about the CLI output:

The portal (Azure's website, 22.17) has a Logs blade (a blade is a page of the portal) that runs the same queries with a chart button; the CLI is what you script, alert-test and paste into incident channels.

Timestamps are UTC

TimeGenerated is UTC (Coordinated Universal Time, the same clock everywhere in the world, no summer time), always. The portal can display local time; the data is UTC. Write incident timelines in UTC and you will never be an hour off during summer time again.

Cost, briefly (the cost lesson has the rest)

Workspaces bill by GB ingested - ingestion is data arriving in the workspace, so you pay for every gigabyte sent in (and for retention past the free period). The Usage table records GB per table per day, and every record carries _BilledSize. A single deployment left at debug log level can multiply the bill - there is a mission about exactly that.

A tour of the Container Insights tables

KubePodInventory   one row per pod per minute: Namespace, Name, PodStatus, Computer (node),
                   PodRestartCount (cumulative!), ContainerLastStatus, ControllerName
KubeNodeInventory  one row per node per minute: Status (Ready/NotReady), Labels (JSON), KubeletVersion
KubeEvents         Kubernetes events: Reason (BackOff, FailedMount, FailedScheduling...), Message, ObjectKind
ContainerLogV2     every stdout/stderr line: PodName, PodNamespace, ContainerName, LogMessage (dynamic), LogLevel
Perf               CPU/memory per node and container: ObjectName, CounterName, CounterValue
Heartbeat          one row per agent per minute - "is this node still reporting?"

Some words in that table: PodRestartCount is how many times the pod's container has been restarted (Ch 15), and cumulative means it only ever grows - it is a running total, not "restarts this minute". dynamic is KQL's type for a JSON value (lesson 24.11).

Inventory tables are snapshots, not event streams: a pod that ran all hour has sixty rows, one per minute. That shapes every query you write against them: you will want the latest row per thing, the growth of a counter (max - min), first and last seen (the smallest and largest TimeGenerated). The next lessons teach the operators for each.

Ingestion latency

Ingestion latency is the delay between something happening and its record being queryable. Records typically arrive 1-3 minutes after the event (more under load). During an incident the last couple of minutes are incomplete; an alert (0.21) that looks back only 1 minute over a table that is 3 minutes late misses things. The column _TimeReceived (real workspaces; not simulated here) shows the delay.

Access

Querying is an RBAC action like any other (22.13): it needs Microsoft.OperationalInsights/workspaces/query/*, which the built-in role Log Analytics Reader on the workspace grants. With resource-context access, a team with Reader on its own AKS cluster can query only the rows that came from that cluster, without rights on the whole workspace - the usual way to share one workspace safely.

What you can now do

Why it helps

During an AKS incident the first question is "where do I look?", and the answer is almost always the workspace: ContainerLogV2 for application output, KubeEvents for evictions and failed mounts, KubePodInventory for restarts. Knowing the tables and that inventory tables are per-minute snapshots saves you from wrong counts and wrong conclusions.

The other lesson is to set things up before the incident. When security asks "who read that secret last Tuesday?", the answer depends on whether a diagnostic setting existed last Tuesday. You'll also run queries from the CLI and paste results into incident channels, and you need to know it wants the customerId, returns strings, and that timestamps are UTC, so your timeline isn't an hour off.

Commands in this lesson

az

FAQ

Why is my Key Vault access not in the logs?

Because resource logs aren't collected by default. Only the activity log, control-plane operations, is recorded automatically. For Key Vault's data-plane access, secret reads and writes, you need a diagnostic setting sending the AuditEvent category to a Log Analytics workspace, storage account or event hub. Without it, there's no record of who read which secret. Enforce diagnostic settings with Azure Policy so every new resource has them from the start.

What's the difference between AzureDiagnostics and resource-specific tables?

AzureDiagnostics is the older shared table where many services write their resource logs, with columns suffixed by type like httpStatusCode_d or identity_claim_oid_g, and it gets very wide and awkward to query. Resource-specific tables, like AKSAudit or AGWAccessLogs, have one schema per log type with proper column names and types. Prefer resource-specific mode in diagnostic settings for services that support it.

Why does the CLI query fail with the workspace name?

az monitor log-analytics query -w wants the workspace's customerId, a GUID, not its name or ARM resource ID. Passing the name fails with PathNotFoundError. Get it with az monitor log-analytics workspace show -g <rg> -n <name> --query customerId -o tsv. Also note that the CLI returns every value as a string and adds a TableName column, which matters when you process results with --query or jq.

Why do inventory tables have so many rows?

KubePodInventory and KubeNodeInventory are snapshots taken every minute, not event logs. A pod that ran for an hour has about sixty rows. So counting rows doesn't count pods, and PodRestartCount is cumulative. Use arg_max(TimeGenerated, *) by Name for the current state, max - min of a counter for growth over a window, and min and max of TimeGenerated for first and last seen.

How fresh is the data in Log Analytics?

Records usually arrive one to three minutes after the event, sometimes more under load. During an incident the last couple of minutes are incomplete, so a spike may look like it's ending when it isn't. For alerts, a lookback shorter than the ingestion delay misses records; give log alerts a window of at least five minutes. Metrics are faster, which is one reason fast alerts are usually metric alerts.

In an interview Mid

What is Azure Monitor, and where do the logs for an AKS platform come from?

Azure Monitor is Azure's monitoring umbrella with three stores: Metrics (numeric time series, cheap, 93 days - dashboards and alerts on known signals), Logs (a Log Analytics workspace: tables queried with KQL - where you go during an incident) and the Activity log (control-plane "who did what", 90 days).

Nothing arrives by itself. Sources:

Usually one workspace per environment, timestamps in UTC, billed per GB ingested, and teams get resource-context access to only their cluster's rows.

Also asked: How do you give application teams access to their logs in a shared workspace without exposing everyone's data? · What would you configure on day one of a new AKS platform so that future incidents can be investigated? · Why are the most recent minutes of logs incomplete during an incident?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.