OnCallReady

Lesson 23.22 · Azure II: Networking & AKS · 14 min read

AKS networking: node subnet, overlay, and the IP bill

In plain words

Imagine a cinema that sells seats in blocks. In one cinema, every time a new row of chairs is added, it immediately reserves thirty seat numbers for that row, whether anyone sits there or not, and the building only has so many numbers. In another cinema, each row gets just one number on the outside, and the seats inside are numbered with a separate private system that nobody outside needs to know.

Those are the two main AKS networking modes. Azure CNI with node subnet reserves maxPods plus one VNet IPs per node the moment the node is created, so a /24 with maxPods 30 fits only eight nodes. Azure CNI Overlay gives each node one VNet IP and numbers pods from a private CIDR like 10.244.0.0/16, NATed to the node when they leave the cluster.

Every pod needs an IP address (16.35). On AKS, where that address comes from decides how big your subnet must be - and whether a Black Friday scale-up works or fails with "not enough IPs". This lesson is the four ways AKS gives pods addresses, and the arithmetic for each.

What you need to know already: 16.35 (CNI: how a pod gets its IP, pod and service CIDRs), 23.1 (subnet sizes, 5 reserved), 23.20 (node pools, surge), 8.9 (Azure subnets and AKS sizing), 23.3 (NSGs).

The Notion question: size a subnet for an AKS cluster on Azure CNI. First, which Azure CNI - the answer is very different.

The options

Azure CNI is Azure's CNI plugin (16.35); it runs in one of three modes. kubenet is the older, simpler plugin. maxPods is the most pods one node may run - a per-pool setting fixed at creation.

modepod IPs come fromdefault maxPodsVNet IPs per node
Azure CNI Overlay (default when --network-plugin is omitted)a private CIDR (default 10.244.0.0/16), /24 per node2501
Azure CNI node subnet ("legacy": --network-plugin azure without a mode)the node subnet, pre-allocated30maxPods + 1
Azure CNI pod subnet (dynamic allocation)a separate pod subnet, allocated in blocks1101 in the node subnet
kubeneta private CIDR, routed with a route table1101 (being retired - do not start new clusters on it)

Pods on node-subnet CNI are first-class VNet citizens: reachable from peered VNets and on-prem by their pod IP, visible to NSGs. Overlay pods are NAT-ed (their source address rewritten, like SNAT in 23.16) to the node IP when they leave the cluster - which is what almost everyone wants, and why overlay is the default.

The node-subnet formula

In node-subnet mode each node reserves its own IP plus maxPods pod IPs at the moment it is created (pre-allocated), whether or not the pods exist. From Microsoft's IP planning guide:

IPs needed = (nodes + surge) + (nodes + surge) * maxPods
           = (nodes + surge) * (maxPods + 1)

then add Azure's 5 reserved and anything else in the subnet (internal load balancer frontends, private endpoints if you put them there).

Worked, the chapter 8 question: 50 nodes, 30 pods each, surge 1:

(50 + 1) * (30 + 1) = 1,581  + 5 reserved = 1,586
/22 = 1,024 - 5 = 1,019   too small
/21 = 2,048 - 5 = 2,043   fits, with room for ~14 more nodes

And this lab's cluster: snet-aks is a /24 = 251 usable, maxPods 30:

251 / 31 = 8.09  ->  8 nodes, total, across ALL pools in the subnet

system 2 + user 3 = 5 nodes = 155 addresses. Scale user to 7 and AKS refuses before creating anything:

ERROR: (InsufficientSubnetSize) Pre-allocated IPs 279 exceeds IPs available 251 in Subnet Cidr
10.20.0.0/24, Subnet Name snet-aks. http://aka.ms/aks/insufficientsubnetsize

279 = (2 + 7) * 31. The error message is the formula.

The trap inside the trap: filling the subnet to exactly 8 nodes leaves no room for a surge node, so the next upgrade fails. Always keep maxSurge worth of nodes free.

Getting out of a full subnet

You cannot resize a subnet in use, and you cannot change a pool's subnet or maxPods. So:

  1. create a bigger subnet (or move to overlay on a new cluster)
  2. add a new node pool in it (--vnet-subnet-id, and a lower --max-pods if node count matters more than density)
  3. cordon and drain the old pool, delete it

For a cluster that is simply growing, the long-term answer is Overlay: node count is limited by the node subnet only (one IP per node), and pod count by the pod CIDR (/24 per node = 256 per node, 250 usable pods).

Overlay

az aks create: create a cluster; --network-plugin azure --network-plugin-mode overlay picks Azure CNI Overlay; --pod-cidr the private range pods get addresses from; --vnet-subnet-id the subnet the nodes go in (the subnet's full resource ID).

az aks create -g rg -n aks-new --network-plugin azure --network-plugin-mode overlay \
  --pod-cidr 10.244.0.0/16 --vnet-subnet-id <node subnet id>

Pod CIDR /16 = 256 nodes worth of /24s. Plan it not to overlap anything the pods talk to directly - but it can be reused across clusters, since it never leaves the cluster.

Service CIDR

--service-cidr (default 10.0.0.0/16: where ClusterIPs come from, 16.1) and --dns-service-ip (default 10.0.0.10: CoreDNS's ClusterIP) are cluster-internal virtual IPs. They must not overlap the VNet or anything routed to it. 10.0.0.0/16 overlaps a lot of corporate networks - pick something deliberate.

maxPods: density vs IPs

Raising maxPods (up to 250) packs more pods per node - fine on overlay, expensive on node subnet (each node reserves maxPods+1). Lowering it below 30 saves IPs but system pods alone take ~10-15 per node. The minimum is 10, and only if the pool still has room for 30 pods in total.

The pod subnet option

Azure CNI pod subnet (--pod-subnet-id) puts pods on routable VNet IPs, but from a separate subnet, allocated to nodes in blocks of 16 as needed instead of pre-allocating maxPods per node:

az aks create -g rg -n aks-ps --network-plugin azure \
  --vnet-subnet-id <node subnet> --pod-subnet-id <pod subnet> --max-pods 110

Pods keep real VNet IPs (on-prem can reach them, NSGs see them) without the node-subnet waste. Size the pod subnet for peak pods, not nodes x maxPods.

Checking a cluster's plan in one query

$ az aks show -g rg-oncall-lab -n aks-sysop --query "{plugin:networkProfile.networkPlugin, mode:networkProfile.networkPluginMode, pods:networkProfile.podCidr, svc:networkProfile.serviceCidr, pools:agentPoolProfiles[].{n:name, count:count, max:maxCount, maxPods:maxPods, subnet:vnetSubnetId}}"

From that you can compute the worst case: for every pool in a subnet, (maxCount or count) + surge, times maxPods + 1 on node-subnet CNI. (maxCount is the cluster autoscaler's upper limit, next lesson.)

What you can now do

Why it helps

Choosing the CNI mode and sizing the subnet are among the few AKS decisions you can't undo without rebuilding. Get them wrong and the cluster can't scale or upgrade: InsufficientSubnetSize with the formula in the error message. "Size a subnet for an AKS cluster" is a standard interview and design review question, and you'll be expected to do the arithmetic on the spot.

The trade-off also matters for security and connectivity: node-subnet pods are first-class VNet addresses, reachable from on-prem and visible to NSGs and firewalls, while overlay pods appear as their node's IP. That changes firewall rules, on-prem allowlists and troubleshooting. And the service CIDR default, 10.0.0.0/16, overlaps many corporate networks, which is worth catching in review.

Commands in this lesson

az

FAQ

How do I size a subnet for Azure CNI node subnet?

Each node takes one IP for itself plus maxPods IPs for pods, pre-allocated when it's created. So IPs needed = (nodes + surge) × (maxPods + 1), plus 5 reserved, plus anything else in the subnet. For 50 nodes, 30 pods each, surge 1: 51 × 31 = 1,581, plus 5 is 1,586, which needs a /21; a /22 has only 1,019 usable. Use the autoscaler's max count for every pool in the subnet.

Why is Overlay the default now?

Because it uses far fewer VNet addresses: one per node, with pods numbered from a private pod CIDR that never leaves the cluster and can be reused across clusters. That removes the main scaling limit of node-subnet CNI and fits most workloads, since pods reach external services through their node's IP. It supports up to 250 pods per node. You choose node subnet or pod subnet mode only when pods need routable VNet IPs.

What is the pod subnet mode?

Azure CNI with a separate pod subnet and dynamic IP allocation. Pods get real VNet IPs, so on-prem can reach them directly and NSGs and firewalls see them, but from a dedicated pod subnet, allocated to nodes in blocks of 16 as needed rather than pre-allocating maxPods per node. You size the pod subnet for peak pods, not nodes × maxPods. It's the middle ground between overlay and node subnet.

Why does the service CIDR matter?

Service IPs are virtual addresses inside the cluster, handled by kube-proxy, and never exist on the VNet. But if the service CIDR overlaps a range pods need to reach, such as the VNet, a peered spoke or on-prem, traffic to those real addresses is captured as service traffic and goes nowhere. The default 10.0.0.0/16 overlaps many corporate networks, so choose --service-cidr and --dns-service-ip deliberately at creation; they can't be changed later.

What does maxPods trade off?

Density against IP addresses. Higher maxPods, up to 250, packs more pods per node, which is efficient on overlay but very expensive on node-subnet CNI, where each node reserves maxPods plus one addresses. Lower maxPods saves addresses, but system pods alone take around 10-15 per node, and the minimum is 10. It's fixed per node pool at creation, so changing it means a new pool.

In an interview Mid

How would you size a subnet for an AKS cluster using Azure CNI?

First ask which Azure CNI:

IPs = (nodes + surge) x (maxPods + 1)  + 5 reserved
50 nodes, maxPods 30, surge 1:  51 x 31 = 1,581 + 5 = 1,586
/22 = 1,019 usable  too small      /21 = 2,043 usable  fits

Count every pool in the subnet at its max (autoscaler maxCount), plus surge for upgrades - a subnet filled to the last node makes the next upgrade fail. Too small fails up front with InsufficientSubnetSize, and the message is the formula.

You cannot resize a used subnet or change a pool's subnet or maxPods, so the way out is a new subnet, a new pool, drain the old one - or Overlay on a new cluster. Also: the service CIDR (default 10.0.0.0/16) must not overlap your networks.

Also asked: Compare Azure CNI Overlay and Azure CNI with a node subnet. When would you choose each? · What happens when an AKS cluster runs out of IP addresses in its subnet? · Why must the cluster's service CIDR not overlap your corporate networks?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.