Traffic doubles at 9 a.m. and halves at 6 p.m. Paying for the peak all night wastes money; sizing for the night drops requests at 9. And batch jobs could run on Azure's leftover capacity at a tenth of the price - if you accept that the machines can vanish. This lesson is the cluster autoscaler and spot pools.
What you need to know already: 17.1 (requests), 17.28 (the HorizontalPodAutoscaler), 17.13 (taints and tolerations), 17.11 (node affinity), 17.15 (topology spread), 17.26 (PDBs), 23.20 (node pools), 23.22 (maxPods, subnet sizing).
The cluster autoscaler
The cluster autoscaler adds nodes when pods are Pending because nothing fits (unschedulable), and removes nodes that have been underused for a while and whose pods fit elsewhere. It does not look at CPU usage - that is the Horizontal Pod Autoscaler's job, which adds pods. HPA adds pods, pods go Pending, the cluster autoscaler adds nodes.
az aks nodepool update changes an existing pool: --enable-cluster-autoscaler --min-count --max-count switches the autoscaler on with its limits, --update-cluster-autoscaler changes the limits, --disable-cluster-autoscaler hands the count back to you.
az aks nodepool update -g rg-oncall-lab --cluster-name aks-sysop -n user \
--enable-cluster-autoscaler --min-count 2 --max-count 5
az aks nodepool update ... --update-cluster-autoscaler --min-count 2 --max-count 6
az aks nodepool update ... --disable-cluster-autoscaler
Once it is on, you stop setting the count - the autoscaler owns it:
# on a pool with the cluster autoscaler enabled
az aks nodepool scale -g rg-oncall-lab --cluster-name aks-sysop -n user -c 4
ERROR: Cannot scale cluster autoscaler enabled node pool.
On a cluster with several pools, az aks update --enable-cluster-autoscaler refuses too ("Please use "az aks nodepool" command to update per node pool auto scaler settings") - it is a per-pool setting.
What it respects, and what blocks scale-down:
- Requests, not usage. Pods without resource requests confuse it completely.
- PodDisruptionBudgets when draining a node.
- Pods with local storage (
emptyDir), without a controller, or annotatedcluster-autoscaler.kubernetes.io/safe-to-evict: "false"pin their node. - max-count x (maxPods+1) must fit the subnet on node-subnet CNI - the autoscaler hits InsufficientSubnetSize just like you did.
The cluster-wide autoscaler profile (settings for all pools) tunes the timing:
$ az aks show -g rg-oncall-lab -n aks-sysop --query autoScalerProfile
{
"balance-similar-node-groups": "false",
"expander": "random",
"max-graceful-termination-sec": "600",
"scale-down-delay-after-add": "10m",
"scale-down-unneeded-time": "10m",
"scan-interval": "10s"
}
Scale-up takes the time to boot a VM (a few minutes); scale-down waits ~10 minutes of being unneeded. Traffic spikes shorter than a node boot are the HPA's and the headroom's job, not the autoscaler's.
Spot node pools
Spot VMs use Azure's spare capacity at a discount of up to ~90%, and Azure can evict them (take them away) at any time - when it needs the capacity back, or when the price rises above your max.
az aks nodepool add creates a pool; the spot-specific flags are explained right below. The regular priority is called on-demand (full price, never evicted).
az aks nodepool add -g rg-oncall-lab --cluster-name aks-sysop -n spot \
--priority Spot --eviction-policy Delete --spot-max-price -1 \
--enable-cluster-autoscaler --min-count 0 --max-count 3 --node-vm-size Standard_D8ds_v5
--spot-max-price -1: pay up to the on-demand price, so you are only ever evicted for capacity, never for price.--eviction-policy Delete: evicted VMs are deleted (the default and the one you want; Deallocate keeps paying for disks).- AKS adds the taint
kubernetes.azure.com/scalesetpriority=spot:NoScheduleand the labelkubernetes.azure.com/scalesetpriority: spotautomatically. Nothing lands on spot unless it tolerates that taint. - Spot pools cannot be System pools.
- Eviction gives ~30 seconds notice (a Preempt scheduled event, which AKS's node problem detector - an agent that watches node health - turns into a
PreemptSchedulednode condition/event).
spec:
tolerations:
- key: kubernetes.azure.com/scalesetpriority
operator: Equal
value: spot
effect: NoSchedule
affinity:
nodeAffinity:
preferredDuringSchedulingIgnoredDuringExecution:
- weight: 100
preference:
matchExpressions:
- key: kubernetes.azure.com/scalesetpriority
operator: In
values: ["spot"]
Tolerate spot and prefer it, but allow regular nodes as the fallback - and spread replicas (topologySpreadConstraints) so one eviction wave cannot take them all.
What belongs on spot: batch jobs, build machines, stateless workers that can be interrupted, dev environments. What does not: anything where an eviction of a majority of replicas is an outage.
Later (Ch 24): the chapter's big incident (24.23) is exactly this mistake, hunted from the logs.
Designing pools
A reasonable default for a production AKS:
system 3 x D4ds_v5 across zones, CriticalAddonsOnly taint, no autoscaler or 3..5
user D8ds_v5, autoscaler 3..N, zones 1 2 3
spot D8ds_v5 spot, autoscaler 0..M, for interruptible work
(+ special pools: memory-optimised, GPU, Windows - each tainted for its workloads)
What the autoscaler says when it will not scale
The autoscaler writes events on the Pending pods:
Normal TriggeredScaleUp pod triggered scale-up: [{aks-user-31415926-vmss 3->5 (max: 6)}]
Normal NotTriggerScaleUp pod didn't trigger scale-up: 1 max node group size reached
Normal NotTriggerScaleUp pod didn't trigger scale-up: 2 node(s) had untolerated taint {sku: gpu}, 1 Insufficient memory
Warning FailedScaleUp Node scale up in zones westeurope-1 associated with this pod failed: ... InsufficientSubnetSize
Read the reason: max reached, nothing could ever fit (taint, selector, size), or Azure refused (subnet, quota - a per-subscription limit on how many vCPUs of a VM family you may run - or spot capacity). You read them with kubectl describe pod or kubectl get events (Ch 15).
Priority expander, briefly
With several pools that could host a pending pod, the expander in the autoscaler profile chooses: random (default), least-waste, priority (a ConfigMap ranking pools - "try spot first, then regular"). The priority expander plus a spot pool and a regular pool with the same shape is the standard "cheap when possible, available always" setup.
What you can now do
- Hand a pool's node count to the cluster autoscaler, and say what blocks its scale-down.
- Add a spot pool and write the toleration and affinity that use it safely.
- Read the autoscaler's events when it refuses to scale.