Every pod needs an IP address (16.35). On AKS, where that address comes from decides how big your subnet must be - and whether a Black Friday scale-up works or fails with "not enough IPs". This lesson is the four ways AKS gives pods addresses, and the arithmetic for each.
What you need to know already: 16.35 (CNI: how a pod gets its IP, pod and service CIDRs), 23.1 (subnet sizes, 5 reserved), 23.20 (node pools, surge), 8.9 (Azure subnets and AKS sizing), 23.3 (NSGs).
The Notion question: size a subnet for an AKS cluster on Azure CNI. First, which Azure CNI - the answer is very different.
The options
Azure CNI is Azure's CNI plugin (16.35); it runs in one of three modes. kubenet is the older, simpler plugin. maxPods is the most pods one node may run - a per-pool setting fixed at creation.
| mode | pod IPs come from | default maxPods | VNet IPs per node |
|---|---|---|---|
Azure CNI Overlay (default when --network-plugin is omitted) | a private CIDR (default 10.244.0.0/16), /24 per node | 250 | 1 |
Azure CNI node subnet ("legacy": --network-plugin azure without a mode) | the node subnet, pre-allocated | 30 | maxPods + 1 |
| Azure CNI pod subnet (dynamic allocation) | a separate pod subnet, allocated in blocks | 110 | 1 in the node subnet |
| kubenet | a private CIDR, routed with a route table | 110 | 1 (being retired - do not start new clusters on it) |
Pods on node-subnet CNI are first-class VNet citizens: reachable from peered VNets and on-prem by their pod IP, visible to NSGs. Overlay pods are NAT-ed (their source address rewritten, like SNAT in 23.16) to the node IP when they leave the cluster - which is what almost everyone wants, and why overlay is the default.
The node-subnet formula
In node-subnet mode each node reserves its own IP plus maxPods pod IPs at the moment it is created (pre-allocated), whether or not the pods exist. From Microsoft's IP planning guide:
IPs needed = (nodes + surge) + (nodes + surge) * maxPods
= (nodes + surge) * (maxPods + 1)
then add Azure's 5 reserved and anything else in the subnet (internal load balancer frontends, private endpoints if you put them there).
Worked, the chapter 8 question: 50 nodes, 30 pods each, surge 1:
(50 + 1) * (30 + 1) = 1,581 + 5 reserved = 1,586
/22 = 1,024 - 5 = 1,019 too small
/21 = 2,048 - 5 = 2,043 fits, with room for ~14 more nodes
And this lab's cluster: snet-aks is a /24 = 251 usable, maxPods 30:
251 / 31 = 8.09 -> 8 nodes, total, across ALL pools in the subnet
system 2 + user 3 = 5 nodes = 155 addresses. Scale user to 7 and AKS refuses before creating anything:
ERROR: (InsufficientSubnetSize) Pre-allocated IPs 279 exceeds IPs available 251 in Subnet Cidr
10.20.0.0/24, Subnet Name snet-aks. http://aka.ms/aks/insufficientsubnetsize
279 = (2 + 7) * 31. The error message is the formula.
The trap inside the trap: filling the subnet to exactly 8 nodes leaves no room for a surge node, so the next upgrade fails. Always keep maxSurge worth of nodes free.
Getting out of a full subnet
You cannot resize a subnet in use, and you cannot change a pool's subnet or maxPods. So:
- create a bigger subnet (or move to overlay on a new cluster)
- add a new node pool in it (
--vnet-subnet-id, and a lower--max-podsif node count matters more than density) - cordon and drain the old pool, delete it
For a cluster that is simply growing, the long-term answer is Overlay: node count is limited by the node subnet only (one IP per node), and pod count by the pod CIDR (/24 per node = 256 per node, 250 usable pods).
Overlay
az aks create: create a cluster; --network-plugin azure --network-plugin-mode overlay picks Azure CNI Overlay; --pod-cidr the private range pods get addresses from; --vnet-subnet-id the subnet the nodes go in (the subnet's full resource ID).
az aks create -g rg -n aks-new --network-plugin azure --network-plugin-mode overlay \
--pod-cidr 10.244.0.0/16 --vnet-subnet-id <node subnet id>
Pod CIDR /16 = 256 nodes worth of /24s. Plan it not to overlap anything the pods talk to directly - but it can be reused across clusters, since it never leaves the cluster.
Service CIDR
--service-cidr (default 10.0.0.0/16: where ClusterIPs come from, 16.1) and --dns-service-ip (default 10.0.0.10: CoreDNS's ClusterIP) are cluster-internal virtual IPs. They must not overlap the VNet or anything routed to it. 10.0.0.0/16 overlaps a lot of corporate networks - pick something deliberate.
maxPods: density vs IPs
Raising maxPods (up to 250) packs more pods per node - fine on overlay, expensive on node subnet (each node reserves maxPods+1). Lowering it below 30 saves IPs but system pods alone take ~10-15 per node. The minimum is 10, and only if the pool still has room for 30 pods in total.
The pod subnet option
Azure CNI pod subnet (--pod-subnet-id) puts pods on routable VNet IPs, but from a separate subnet, allocated to nodes in blocks of 16 as needed instead of pre-allocating maxPods per node:
az aks create -g rg -n aks-ps --network-plugin azure \
--vnet-subnet-id <node subnet> --pod-subnet-id <pod subnet> --max-pods 110
Pods keep real VNet IPs (on-prem can reach them, NSGs see them) without the node-subnet waste. Size the pod subnet for peak pods, not nodes x maxPods.
Checking a cluster's plan in one query
$ az aks show -g rg-oncall-lab -n aks-sysop --query "{plugin:networkProfile.networkPlugin, mode:networkProfile.networkPluginMode, pods:networkProfile.podCidr, svc:networkProfile.serviceCidr, pools:agentPoolProfiles[].{n:name, count:count, max:maxCount, maxPods:maxPods, subnet:vnetSubnetId}}"
From that you can compute the worst case: for every pool in a subnet, (maxCount or count) + surge, times maxPods + 1 on node-subnet CNI. (maxCount is the cluster autoscaler's upper limit, next lesson.)
What you can now do
- Name the four ways AKS gives pods IPs and what each costs in VNet addresses.
- Size a node subnet:
(nodes + surge) x (maxPods + 1)plus 5, rounded up to a prefix. - Explain why a full subnet means a new subnet and a new pool, and why Overlay avoids it.