OnCallReady

Chapter 24 Azure III: Monitor, KQL & Cost

Log Analytics and KQL properly - summarize, joins, parsing, the incident playbook - two outages hunted from logs, and where the money goes.

In plain words

Think of a detective who arrives after something went wrong in a big building. She doesn't guess: she reads the logbooks. The door log says who came in and when, the lift log says which floors were visited, the kitchen log says what was cooked. None of them alone tells the story, but lined up by time, they do. And at the end of the month, the building manager asks a different question with the same logbooks: why is the electricity bill so high?

Azure Monitor keeps those logbooks in a Log Analytics workspace, and KQL is the detective's method for reading them: filter by time, summarise, join tables, pull fields out of text. This chapter teaches that method with ContainerLogV2, KubePodInventory, KubeEvents and AppRequests, then turns it to cost: where the money in an AKS platform really goes.

Why it matters on call

When production breaks on AKS, the evidence lives in Log Analytics, not on a box you can SSH into: container logs, pod and node snapshots, Kubernetes events, Key Vault audits, App Gateway access logs, the activity log. Being the person who can answer "what broke, since when, what changed, who's affected" with six queries in the first ten minutes is one of the most valued skills on an SRE team.

KQL also runs the alerts, where the integer-division bug makes them never fire, and the cost analysis, where one debug-level deployment can outspend a node pool. This chapter comes after Azure identity and networking because it reads the logs those resources produce: Key Vault audits, App Gateway access logs, AKS events. The query habits you build here (time first, summarize, last seen) carry over to any monitoring tool you use later.

Lessons

  1. Azure Monitor: where the logs actually are
  2. KQL 1: pipes, filters and strings
  3. KQL 2: summarize, bin, percentiles
  4. KQL 3: join, union and let for tables
  5. KQL 4: getting fields out of text and JSON
  6. KQL 5: the incident playbook
  7. Cost: where the money goes in an AKS platform

21 hands-on labs (missions, incidents and drills) run in the terminal: Open this chapter in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.

Questions people ask

Why not just use kubectl logs?

kubectl logs shows one pod's recent output, and only while the pod exists: a pod that crashed and was replaced, or evicted with its node, takes its logs with it. It also cannot answer questions across pods ("which of 40 pods returned 503s in the last hour?") or join logs with anything else. Log Analytics keeps every container's output for the retention period, next to pod and node snapshots, Kubernetes events, Key Vault audits and the activity log, so one query can line them up by time. Use kubectl logs for a quick look at a live pod; use the workspace for incidents and anything after the fact.

What is the difference between metrics and logs in Azure Monitor?

Metrics are numeric time series at one-minute granularity, kept for 93 days, cheap and fast, ideal for dashboards and alerts on known signals like CPU or request counts. Logs are records in tables in a Log Analytics workspace, queried with KQL, where you can filter, join and aggregate any way you like. You pay per GB ingested. Logs are where you go during an incident, because you can ask questions nobody planned for.

Is KQL like SQL?

It covers the same ground, filtering, projecting, grouping, joining, but reads as a pipeline from left to right: a table, then operators separated by |, each taking a table and returning one. where is WHERE, summarize ... by is GROUP BY, project is SELECT, join is JOIN. It's closer to a shell pipeline than to SQL's nested clauses, which makes building a query step by step natural, and it has strong time-series functions like bin.

Why is cost in the same chapter as monitoring?

Because they're analysed with the same tools and they're linked. The Usage table and _BilledSize column are queried with KQL, and log ingestion itself is often the second-biggest line item after compute. Understanding cost also needs the platform knowledge of the earlier chapters: requests and the autoscaler for compute, reclaim policies for orphaned disks, SNAT and public IPs for networking. It's operational work, not finance's job alone.

How realistic are the lab's tables?

The table and column names, types and shapes match real Container Insights, Application Insights and diagnostic tables, and the KQL interpreter implements the operators taught: where, project, extend, summarize, bin, join kinds, union, let, parse, extract, parse_json, arg_max and more. A few things are marked as not simulated, such as lookup and where * has across all columns, and _TimeReceived. Queries you write here work in a real workspace.