OnCallReady

Lesson 23.25 · Azure II: Networking & AKS · 14 min read

Cluster autoscaler, spot pools, and pool design

In plain words

Imagine a restaurant that adds tables when guests are left standing at the door with nowhere to sit, and removes tables that stay empty for a while. It doesn't look at how much people eat; it only asks "is someone waiting for a seat?". Some extra tables are borrowed very cheaply from the shop next door, but the neighbour can take them back with thirty seconds' warning.

The cluster autoscaler is that restaurant manager: it adds nodes when pods are Pending because nothing fits, and removes underused nodes whose pods fit elsewhere. It looks at requests, not CPU usage; that's the HPA's job, which adds pods. Spot node pools are the borrowed tables: up to about 90% cheaper, tainted kubernetes.azure.com/scalesetpriority=spot:NoSchedule, and evictable at any time.

Traffic doubles at 9 a.m. and halves at 6 p.m. Paying for the peak all night wastes money; sizing for the night drops requests at 9. And batch jobs could run on Azure's leftover capacity at a tenth of the price - if you accept that the machines can vanish. This lesson is the cluster autoscaler and spot pools.

What you need to know already: 17.1 (requests), 17.28 (the HorizontalPodAutoscaler), 17.13 (taints and tolerations), 17.11 (node affinity), 17.15 (topology spread), 17.26 (PDBs), 23.20 (node pools), 23.22 (maxPods, subnet sizing).

The cluster autoscaler

The cluster autoscaler adds nodes when pods are Pending because nothing fits (unschedulable), and removes nodes that have been underused for a while and whose pods fit elsewhere. It does not look at CPU usage - that is the Horizontal Pod Autoscaler's job, which adds pods. HPA adds pods, pods go Pending, the cluster autoscaler adds nodes.

az aks nodepool update changes an existing pool: --enable-cluster-autoscaler --min-count --max-count switches the autoscaler on with its limits, --update-cluster-autoscaler changes the limits, --disable-cluster-autoscaler hands the count back to you.

az aks nodepool update -g rg-oncall-lab --cluster-name aks-sysop -n user \
  --enable-cluster-autoscaler --min-count 2 --max-count 5
az aks nodepool update ... --update-cluster-autoscaler --min-count 2 --max-count 6
az aks nodepool update ... --disable-cluster-autoscaler

Once it is on, you stop setting the count - the autoscaler owns it:

# on a pool with the cluster autoscaler enabled
az aks nodepool scale -g rg-oncall-lab --cluster-name aks-sysop -n user -c 4
ERROR: Cannot scale cluster autoscaler enabled node pool.

On a cluster with several pools, az aks update --enable-cluster-autoscaler refuses too ("Please use "az aks nodepool" command to update per node pool auto scaler settings") - it is a per-pool setting.

What it respects, and what blocks scale-down:

The cluster-wide autoscaler profile (settings for all pools) tunes the timing:

$ az aks show -g rg-oncall-lab -n aks-sysop --query autoScalerProfile
{
  "balance-similar-node-groups": "false",
  "expander": "random",
  "max-graceful-termination-sec": "600",
  "scale-down-delay-after-add": "10m",
  "scale-down-unneeded-time": "10m",
  "scan-interval": "10s"
}

Scale-up takes the time to boot a VM (a few minutes); scale-down waits ~10 minutes of being unneeded. Traffic spikes shorter than a node boot are the HPA's and the headroom's job, not the autoscaler's.

Spot node pools

Spot VMs use Azure's spare capacity at a discount of up to ~90%, and Azure can evict them (take them away) at any time - when it needs the capacity back, or when the price rises above your max.

az aks nodepool add creates a pool; the spot-specific flags are explained right below. The regular priority is called on-demand (full price, never evicted).

az aks nodepool add -g rg-oncall-lab --cluster-name aks-sysop -n spot \
  --priority Spot --eviction-policy Delete --spot-max-price -1 \
  --enable-cluster-autoscaler --min-count 0 --max-count 3 --node-vm-size Standard_D8ds_v5
spec:
  tolerations:
    - key: kubernetes.azure.com/scalesetpriority
      operator: Equal
      value: spot
      effect: NoSchedule
  affinity:
    nodeAffinity:
      preferredDuringSchedulingIgnoredDuringExecution:
        - weight: 100
          preference:
            matchExpressions:
              - key: kubernetes.azure.com/scalesetpriority
                operator: In
                values: ["spot"]

Tolerate spot and prefer it, but allow regular nodes as the fallback - and spread replicas (topologySpreadConstraints) so one eviction wave cannot take them all.

What belongs on spot: batch jobs, build machines, stateless workers that can be interrupted, dev environments. What does not: anything where an eviction of a majority of replicas is an outage.

Later (Ch 24): the chapter's big incident (24.23) is exactly this mistake, hunted from the logs.

Designing pools

A reasonable default for a production AKS:

system   3 x D4ds_v5 across zones, CriticalAddonsOnly taint, no autoscaler or 3..5
user     D8ds_v5, autoscaler 3..N, zones 1 2 3
spot     D8ds_v5 spot, autoscaler 0..M, for interruptible work
(+ special pools: memory-optimised, GPU, Windows - each tainted for its workloads)

What the autoscaler says when it will not scale

The autoscaler writes events on the Pending pods:

Normal   TriggeredScaleUp   pod triggered scale-up: [{aks-user-31415926-vmss 3->5 (max: 6)}]
Normal   NotTriggerScaleUp  pod didn't trigger scale-up: 1 max node group size reached
Normal   NotTriggerScaleUp  pod didn't trigger scale-up: 2 node(s) had untolerated taint {sku: gpu}, 1 Insufficient memory
Warning  FailedScaleUp      Node scale up in zones westeurope-1 associated with this pod failed: ... InsufficientSubnetSize

Read the reason: max reached, nothing could ever fit (taint, selector, size), or Azure refused (subnet, quota - a per-subscription limit on how many vCPUs of a VM family you may run - or spot capacity). You read them with kubectl describe pod or kubectl get events (Ch 15).

Priority expander, briefly

With several pools that could host a pending pod, the expander in the autoscaler profile chooses: random (default), least-waste, priority (a ConfigMap ranking pools - "try spot first, then regular"). The priority expander plus a spot pool and a regular pool with the same shape is the standard "cheap when possible, available always" setup.

What you can now do

Why it helps

Autoscaling decides your cluster's cost and whether traffic spikes cause outages. You'll debug "pods stuck Pending and no new nodes" by reading the autoscaler's events: max reached, a taint or selector nothing matches, or Azure refusing because of subnet, quota or spot capacity. And you'll explain why pods without resource requests confuse it completely.

Spot pools are where the savings are, and where one of the more painful incidents comes from: putting a majority of a service's replicas on spot, and an eviction wave taking it down. Knowing to tolerate but only prefer spot, spread replicas and keep a regular fallback pool is exactly the kind of judgement a platform team is expected to encode in its standards.

Commands in this lesson

az

FAQ

What is the difference between the cluster autoscaler and the HPA?

The Horizontal Pod Autoscaler changes the number of pod replicas based on metrics like CPU usage or custom metrics. The cluster autoscaler changes the number of nodes based on whether pods can be scheduled: it adds nodes when pods are Pending because no node has room for their requests, and removes nodes that are underused. They work together: HPA adds pods, pods go Pending, the cluster autoscaler adds nodes.

Why won't the autoscaler scale down a node?

Common reasons: pods on it can't move elsewhere because of their requests or affinity; a PodDisruptionBudget doesn't allow evicting them; pods use local storage such as emptyDir; pods aren't managed by a controller; pods are annotated cluster-autoscaler.kubernetes.io/safe-to-evict: "false"; kube-system pods without PDBs; or the node hasn't been unneeded for long enough, 10 minutes by default. The pool's minimum count also stops it.

Why can't I scale an autoscaler-enabled pool by hand?

Because once the cluster autoscaler is on, it owns the node count within min and max, and a manual count would conflict with it. az aks nodepool scale refuses with Cannot scale cluster autoscaler enabled node pool. Change the min and max with --update-cluster-autoscaler instead, or disable the autoscaler for that pool. It's a per-pool setting, which is why the cluster-level command refuses on multi-pool clusters.

What belongs on spot nodes?

Work that tolerates interruption: batch jobs, CI runners, queue workers that can retry, data processing that checkpoints, and dev or test environments. Not stateful services, and not a majority of a service's replicas, because Azure can evict many spot VMs at once when it needs capacity, with about 30 seconds' notice. Tolerate the spot taint and prefer spot nodes, but allow regular nodes too, and spread replicas across nodes and zones.

What does spot-max-price -1 mean?

You're willing to pay up to the regular on-demand price, so you'll never be evicted because the spot price rose, only when Azure needs the capacity back. That's the usual choice, since price-based evictions add a failure mode without much saving. Combine it with --eviction-policy Delete, the default, so evicted VMs are deleted instead of deallocated and don't keep charging for disks.

In an interview Mid

Pods are Pending but the cluster autoscaler isn't adding nodes. What do you check?

The cluster autoscaler adds nodes only for pods that are Pending because nothing fits, judged by requests, not usage. It writes its reason as events on the pod: kubectl describe pod / kubectl get events.

And timing: a new node takes minutes to boot, so short spikes are the HPA's and headroom's job. Scale-down is blocked by PDBs, emptyDir pods, bare pods and safe-to-evict: "false".

Also asked: How does the cluster autoscaler decide to add or remove nodes? · How would you use spot node pools without risking availability? · What is the difference between the cluster autoscaler and the Horizontal Pod Autoscaler?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.