In the cloud, every resource you create is billed by the hour or by the gigabyte, and nobody sends you the bill until a month later. A forgotten 1 TB disk or one app logging too much quietly costs more than an engineer's laptop. This lesson is where the money goes in an AKS platform and how to find it.
What you need to know already: 23.20-23.25 (AKS, node pools, VM sizes, spot, the cluster autoscaler), 17.1 (requests and limits), 17.26 (PDBs), 16.39 and 16.44 (PVCs, disks, the Retain reclaim policy), 24.1 (Log Analytics and ingestion), 22.17 (tags), 0.8 (SLOs and SLAs), Ch 7 (awk).
The Notion question: where does the money actually go in an AKS cluster? Almost always, in this order:
1. compute the node VMs (VM scale sets) - typically 60-80% of the bill
2. logs Log Analytics ingestion - the surprise line item
3. storage managed disks for PVCs, snapshots, storage accounts
4. network load balancers, public IPs, NAT gateway, egress, App Gateway / WAF, private endpoints
5. the cluster itself: the Standard tier uptime SLA (per cluster per hour)
Words in that list: a VM scale set (VMSS) is the Azure resource behind a node pool - a group of identical VMs that Azure can grow and shrink (23.20). Egress is data leaving Azure (to the internet or another region), billed per GB. A NAT gateway gives many VMs one shared outgoing public IP. The uptime SLA is Microsoft's promise of how available the AKS control plane will be (an SLA, 0.8); you pay for it on the Standard tier.
The control plane (the Kubernetes API server and friends, Ch 15) is free on the Free tier and a fixed hourly fee on Standard/Premium; it is rarely the problem. Nodes are.
Compute levers
- Right-size requests. The autoscaler and the scheduler work from requests. Requests set to twice the real usage mean twice the nodes. Container Insights'
Perf/InsightsMetricsshow actual usage per container; compare. - Scale down. The cluster autoscaler removes underused nodes (10 minutes unneeded by default) - unless something pins them: pods no controller would recreate (a bare pod, not from a Deployment), pods with local storage, or PDBs (17.26) too strict to let any pod move. A cluster that never scales down is usually one of those.
- Stop dev clusters at night:
az aks stop -g <rg> -n <cluster>stops the nodes (you stop paying for them) andaz aks startbrings them back. - Spot for interruptible work: up to ~90% off, can vanish in 30 seconds (chapter 23, and the spot incident).
- Commitments for the steady baseline (the capacity you know you will use all year) - you promise to pay for a period whether you use it or not, and get a discount for promising: · Reserved instances: 1 or 3 years for a specific VM family and region, the deepest discount (up to ~70%), no flexibility. · Savings plans for compute: commit to an hourly spend for 1 or 3 years, applied across VM families and regions - less discount, much more flexible. Commit to the floor you are sure of (system pools, the minimum of user pools), never to the peak.
The log bill
Log Analytics charges per GB ingested (24.1) on the normal Analytics plan - with a commitment tier discount if you promise at least 100 GB/day - plus retention beyond the included period. Container logs are the usual culprit: one deployment left at DEBUG can outspend a node pool. The tools:
The Usage table has one row per table (DataType) per hour with the volume in MB (Quantity) and whether it is billed (IsBillable). Every record in every table also carries _BilledSize, its size in bytes. So:
Usage
| where TimeGenerated > ago(30d) and IsBillable
| summarize GB = round(sum(Quantity) / 1024, 1) by DataType
| order by GB desc
ContainerLogV2
| where TimeGenerated > ago(1h)
| summarize MB = round(sum(_BilledSize) / 1048576.0, 2) by PodNamespace, ContainerName, LogLevel
| top 10 by MB
Fixes, cheapest first:
- fix the log level at the source (the setting that decides how chatty an app is: DEBUG logs everything, INFO the normal events, ERROR only failures);
- filter at collection: a data collection rule (DCR) is the Azure resource that tells the agent what to collect, and Container Insights' DCR can exclude namespaces and log levels; ingestion-time transformations (a KQL filter Azure applies as data arrives) can drop or trim rows;
- move high-volume, rarely-queried tables to the Basic plan (much cheaper ingestion, limited KQL, short interactive retention -
ContainerLogV2supports it); - set a daily cap (a GB-per-day limit after which the workspace stops accepting data) on dev workspaces only (a cap on prod means losing logs exactly
during the incident that generated them).
Orphans
An orphan is a resource that outlived the thing that used it and bills forever. Three list commands, each with a JMESPath filter (22.4): unattached disks, public IPs attached to nothing, and every snapshot (a point-in-time copy of a disk):
az disk list --query "[?diskState=='Unattached'].{name:name, gb:diskSizeGb, sku:sku.name, rg:resourceGroup}" -o table
az network public-ip list --query "[?ipConfiguration==null].{name:name, rg:resourceGroup}" -o table
az snapshot list --query "[].{name:name, created:timeCreated}" -o table
Unattached disks from deleted PVCs (Retain policy), public IPs from deleted Services, old snapshots, forgotten test clusters. The kubernetes.io-created-for-pvc-name tag tells you where a disk came from before you delete it.
Seeing the bill
az consumption usage list --start-date A --end-date B: the subscription's usage, one row per resource per day, with consumedService (e.g. Microsoft.Compute), instanceName (the resource) and pretaxCost. The pipeline below prints three tab-separated columns, then awk (Ch 7) adds up cost per service and sort -k2 -nr sorts by the second column, numerically, biggest first:
az consumption usage list --start-date 2026-08-24 --end-date 2026-09-23 \
--query "[].{svc:consumedService, name:instanceName, cost:pretaxCost}" -o tsv \
| awk -F'\t' '{s[$1]+=$3} END {for (k in s) printf "%-35s %10.2f\n", k, s[k]}' | sort -k2 -nr
pretaxCost comes back as a string - JMESPath cannot sum strings across items (and cannot group at all), which is why this is -o tsv | awk, chapter 7 style. (Newer tooling: Cost Management exports to a storage account, and the az costmanagement extension - an extension is an optional add-on for az you install separately; the consumption API is the one that works everywhere.)
Budgets: why nobody notices the bill
Engineers do not see invoices; finance sees them a month later, aggregated. By then the debug logging has run for five weeks. So:
A budget is an Azure object that tracks spend against an amount you set and emails people as it gets close. az consumption budget create: --budget-name its name; --amount the limit in the billing currency; --time-grain Monthly the amount resets every month; --start-date/--end-date how long the budget runs; --category Cost track money (not usage).
az consumption budget create --budget-name oncall-lab-monthly --amount 1500 \
--time-grain Monthly --start-date 2026-09-01 --end-date 2027-08-31 --category Cost
A budget with alerts at 50/80/100% (and a forecast alert - one that fires when Azure predicts you will cross the amount by month end) to the team that can act, per subscription or resource group; tags (team, env, cost-center, 22.17) made mandatory by Azure Policy (Azure's rule engine that can refuse to create resources that break a rule) so costs can be shown back per team - showback; and a monthly five-minute look at the top ten line items. Cost is an SLO (0.8) like any other: nobody minds spending, everyone minds surprises.
What you can now do
- Name the five places AKS money goes, biggest first, and the lever for each.
- Find orphaned disks and public IPs with
az ... list --query. - Sum a month's cost per service with
az consumption usage list -o tsv | awk, and set a budget so the next surprise is an email.