OnCallReady

Lesson 24.26 · Azure III: Monitor, KQL & Cost · 14 min read

Cost: where the money goes in an AKS platform

In plain words

Imagine your family's monthly bills. Most of the money goes on rent, the big fixed thing. Then there's a surprise: the phone bill doubled because someone left video streaming on all month. There are also subscriptions nobody remembers signing up for, still charging every month. And nobody notices until the bank statement arrives weeks later.

An AKS platform's bill looks the same. Compute, the node VMs, is the rent: typically 60-80%. Log Analytics ingestion is the surprise, when one deployment at DEBUG outspends a node pool. Orphaned disks and public IPs are the forgotten subscriptions. And because engineers don't see invoices, the fixes are right-sizing requests, spot and commitments for compute, filtering logs at collection, cleaning orphans, and budgets with alerts before the statement arrives.

In the cloud, every resource you create is billed by the hour or by the gigabyte, and nobody sends you the bill until a month later. A forgotten 1 TB disk or one app logging too much quietly costs more than an engineer's laptop. This lesson is where the money goes in an AKS platform and how to find it.

What you need to know already: 23.20-23.25 (AKS, node pools, VM sizes, spot, the cluster autoscaler), 17.1 (requests and limits), 17.26 (PDBs), 16.39 and 16.44 (PVCs, disks, the Retain reclaim policy), 24.1 (Log Analytics and ingestion), 22.17 (tags), 0.8 (SLOs and SLAs), Ch 7 (awk).

The Notion question: where does the money actually go in an AKS cluster? Almost always, in this order:

1. compute     the node VMs (VM scale sets) - typically 60-80% of the bill
2. logs        Log Analytics ingestion - the surprise line item
3. storage     managed disks for PVCs, snapshots, storage accounts
4. network     load balancers, public IPs, NAT gateway, egress, App Gateway / WAF, private endpoints
5. the cluster itself: the Standard tier uptime SLA (per cluster per hour)

Words in that list: a VM scale set (VMSS) is the Azure resource behind a node pool - a group of identical VMs that Azure can grow and shrink (23.20). Egress is data leaving Azure (to the internet or another region), billed per GB. A NAT gateway gives many VMs one shared outgoing public IP. The uptime SLA is Microsoft's promise of how available the AKS control plane will be (an SLA, 0.8); you pay for it on the Standard tier.

The control plane (the Kubernetes API server and friends, Ch 15) is free on the Free tier and a fixed hourly fee on Standard/Premium; it is rarely the problem. Nodes are.

Compute levers

The log bill

Log Analytics charges per GB ingested (24.1) on the normal Analytics plan - with a commitment tier discount if you promise at least 100 GB/day - plus retention beyond the included period. Container logs are the usual culprit: one deployment left at DEBUG can outspend a node pool. The tools:

The Usage table has one row per table (DataType) per hour with the volume in MB (Quantity) and whether it is billed (IsBillable). Every record in every table also carries _BilledSize, its size in bytes. So:

Usage
| where TimeGenerated > ago(30d) and IsBillable
| summarize GB = round(sum(Quantity) / 1024, 1) by DataType
| order by GB desc

ContainerLogV2
| where TimeGenerated > ago(1h)
| summarize MB = round(sum(_BilledSize) / 1048576.0, 2) by PodNamespace, ContainerName, LogLevel
| top 10 by MB

Fixes, cheapest first:

during the incident that generated them).

Orphans

An orphan is a resource that outlived the thing that used it and bills forever. Three list commands, each with a JMESPath filter (22.4): unattached disks, public IPs attached to nothing, and every snapshot (a point-in-time copy of a disk):

az disk list --query "[?diskState=='Unattached'].{name:name, gb:diskSizeGb, sku:sku.name, rg:resourceGroup}" -o table
az network public-ip list --query "[?ipConfiguration==null].{name:name, rg:resourceGroup}" -o table
az snapshot list --query "[].{name:name, created:timeCreated}" -o table

Unattached disks from deleted PVCs (Retain policy), public IPs from deleted Services, old snapshots, forgotten test clusters. The kubernetes.io-created-for-pvc-name tag tells you where a disk came from before you delete it.

Seeing the bill

az consumption usage list --start-date A --end-date B: the subscription's usage, one row per resource per day, with consumedService (e.g. Microsoft.Compute), instanceName (the resource) and pretaxCost. The pipeline below prints three tab-separated columns, then awk (Ch 7) adds up cost per service and sort -k2 -nr sorts by the second column, numerically, biggest first:

az consumption usage list --start-date 2026-08-24 --end-date 2026-09-23 \
  --query "[].{svc:consumedService, name:instanceName, cost:pretaxCost}" -o tsv \
  | awk -F'\t' '{s[$1]+=$3} END {for (k in s) printf "%-35s %10.2f\n", k, s[k]}' | sort -k2 -nr

pretaxCost comes back as a string - JMESPath cannot sum strings across items (and cannot group at all), which is why this is -o tsv | awk, chapter 7 style. (Newer tooling: Cost Management exports to a storage account, and the az costmanagement extension - an extension is an optional add-on for az you install separately; the consumption API is the one that works everywhere.)

Budgets: why nobody notices the bill

Engineers do not see invoices; finance sees them a month later, aggregated. By then the debug logging has run for five weeks. So:

A budget is an Azure object that tracks spend against an amount you set and emails people as it gets close. az consumption budget create: --budget-name its name; --amount the limit in the billing currency; --time-grain Monthly the amount resets every month; --start-date/--end-date how long the budget runs; --category Cost track money (not usage).

az consumption budget create --budget-name oncall-lab-monthly --amount 1500 \
  --time-grain Monthly --start-date 2026-09-01 --end-date 2027-08-31 --category Cost

A budget with alerts at 50/80/100% (and a forecast alert - one that fires when Azure predicts you will cross the amount by month end) to the team that can act, per subscription or resource group; tags (team, env, cost-center, 22.17) made mandatory by Azure Policy (Azure's rule engine that can refuse to create resources that break a rule) so costs can be shown back per team - showback; and a monthly five-minute look at the top ten line items. Cost is an SLO (0.8) like any other: nobody minds spending, everyone minds surprises.

What you can now do

Why it helps

Platform teams are increasingly asked "why does our AKS cost this much?", and FinOps is part of the job at any company running cloud at scale. With this lesson you can answer with data: the Usage table and _BilledSize for logs, az consumption usage list for the bill, unattached disks from Retain policies, and requests compared with real usage for compute.

The levers are concrete and often large: halving inflated requests halves node count, a debug log level fixed at the source can cut the log bill dramatically, and reserved instances or savings plans for the steady baseline are significant savings. You'll also know the traps: a daily cap on the production workspace loses logs during the incident that generated them, and committing to the peak instead of the floor wastes money.

FAQ

Why is compute usually the biggest cost?

Because the node VMs run all the time and are sized by what pods request, not what they use. If requests are set to twice real usage, the scheduler and autoscaler provision twice the nodes. Clusters that never scale down, because of pods without controllers, local storage or strict PDBs, keep paying for idle capacity. Right-sizing requests from real usage in Perf or InsightsMetrics is usually the single biggest lever.

How do I find what's driving the log bill?

Query it. The Usage table records billable GB per table per day: Usage | where IsBillable | summarize GB = sum(Quantity) / 1024 by DataType. Every record also carries _BilledSize, so ContainerLogV2 | summarize MB = sum(_BilledSize) / 1048576.0 by PodNamespace, ContainerName, LogLevel finds the noisy workloads. Usually a few containers at debug level, or chatty health-check logs, account for most of the volume.

Should I set a daily cap on the workspace?

Only on dev and test workspaces. When the cap is reached, ingestion stops until the next day, and in production that means losing logs exactly when something is generating a lot of them, which is usually during an incident. For production, control volume at the source, with data collection rules and ingestion-time transformations, and use the Basic plan for high-volume, rarely queried tables. Alert on unusual ingestion growth instead of capping it.

Reserved instances or savings plans?

Reserved instances commit to a specific VM family and region for one or three years, for the deepest discount, but no flexibility if you change VM sizes or regions. Savings plans for compute commit to an hourly spend for one or three years, applied across VM families and regions, with a smaller discount and much more flexibility. Either way, commit only to the baseline you're sure of, system pools and minimum node counts, never to the peak.

Why do orphaned resources happen?

Because some resources outlive what created them. Disks from PVCs with reclaimPolicy: Retain stay after the PVC is deleted, public IPs remain after a LoadBalancer Service was removed in some cases, snapshots are kept forever, and test clusters get forgotten. They're billed every hour. Regular queries like az disk list --query "[?diskState=='Unattached']", owner tags enforced by policy and Advisor recommendations keep them in check.

In an interview Mid

How would you reduce the cost of an AKS cluster?

First know where it goes: compute (node VMs, 60-80%), logs (Log Analytics ingestion - the surprise), storage, network, the cluster tier.

Compute:

Logs: find the biggest tables with Usage | summarize GB = sum(Quantity) / 1024 by DataType, then fix the log level at the source, filter at collection (DCR), Basic plan for noisy rarely-queried tables, a daily cap on dev only.

Orphans: az disk list --query "[?diskState=='Unattached']", public IPs with no ipConfiguration, old snapshots.

Then keep it fixed: budgets with forecast alerts, mandatory tags for showback per team.

Also asked: What are the main cost components of running AKS? · How do you find orphaned Azure resources that are still being billed? · How do you make teams aware of what their workloads cost?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.