OnCallReady

Lesson 23.1 · Azure II: Networking & AKS · 19 min read

VNets and subnets: the five addresses you never get

In plain words

Imagine a street where every house needs a number. The council keeps a few numbers on every street for itself: number 0 is the street sign, 1 is the post box, 2 and 3 are the phone boxes, and the last number is the notice board. So a street built for eight houses really has room for only three families. And once people have moved in, you can't stretch the street; you have to build a new one.

A subnet is that street. Azure keeps .0, .1, .2, .3 and the last address of every subnet, so usable addresses are 2^(32 - prefix) - 5, and the first one you ever get is .4. A /29 holds exactly three things, as SubnetIsFull shows after the third private endpoint. And a subnet with anything in it can't be resized.

On your VM the network was one interface and one address (Ch 1, Ch 8). In Azure you design the network yourself: which private addresses exist, how they are split up, who may talk to whom. Get the address plan wrong and it cannot be fixed later without moving everything. This chapter is Azure networking, then AKS on top of it.

What you need to know already: 8.3 (CIDR: /24 = 256 addresses), 8.6 (carving an address space, overlap), 8.9 (Azure subnets and AKS sizing), 8.11 (routing tables), 22.1 (subscriptions, resource groups, regions, az), 22.4 (--query), 12.20 (Terraform for_each).

Virtual networks and subnets

A virtual network (VNet) is a private address range (for example 10.20.0.0/16) that you own inside Azure, in one region and one subscription. Only things you put in it get those addresses, and by default nothing outside can reach them. A subnet is a slice of the VNet's range (10.20.4.0/24) - the same idea as in Ch 8.

Everything with a private IP sits in a subnet: a VM's NIC (network interface card - the virtual network port of a VM), AKS nodes (and, with one of the AKS network modes you will meet in 23.22, the pods too), private endpoints (23.10), internal load balancers and App Gateway instances (23.16).

az network vnet show -g rg-oncall-lab -n vnet-sysop --query ...: show one VNet; the JMESPath (22.4) picks its address space and each subnet's name and range.

$ az network vnet show -g rg-oncall-lab -n vnet-sysop --query "{space:addressSpace.addressPrefixes, subnets:subnets[].{name:name, prefix:addressPrefix}}"
{
  "space": [
    "10.20.0.0/16"
  ],
  "subnets": [
    {
      "name": "snet-aks",
      "prefix": "10.20.0.0/24"
    },
    {
      "name": "snet-appgw",
      "prefix": "10.20.1.0/24"
    },
    {
      "name": "snet-pe",
      "prefix": "10.20.2.0/28"
    },
    {
      "name": "snet-app",
      "prefix": "10.20.4.0/24"
    }
  ]
}

Five addresses per subnet belong to Azure

Chapter 8 did the arithmetic (8.9); here it is on a real subnet. In every subnet:

x.x.x.0     network address
x.x.x.1     default gateway (the Azure router)
x.x.x.2     \  Azure DNS
x.x.x.3     /
x.x.x.255   broadcast (the last address, whatever the prefix)

So usable = 2^(32 - prefix) - 5, and the first address you ever get is .4:

prefixaddressesusable
/2983 (.4 .5 .6) - the smallest subnet Azure allows
/281611
/273227
/266459
/24256251
/2210241019

A private endpoint (23.10) takes one address. Here four are created in a loop into a /29 subnet called snet-pe2; each prints the IP it got:

# the /29 mission: four endpoints into snet-pe2
for i in 1 2 3 4; do az network private-endpoint create ... --subnet snet-pe2 --query "customDnsConfigs[0].ipAddresses[0]" -o tsv; done
10.20.3.4
10.20.3.5
10.20.3.6
ERROR: (SubnetIsFull) Subnet snet-pe2 with address prefix 10.20.3.0/29 does not have enough
capacity for 1 IP addresses.

A /29 holds three things. Not eight, not six.

Creating subnets, and the two errors you will meet

az network vnet subnet create -g <rg> --vnet-name <vnet> -n <name> --address-prefixes <cidr>: add a subnet to a VNet; --address-prefixes its range.

az network vnet subnet create -g rg-oncall-lab --vnet-name vnet-sysop -n snet-data --address-prefixes 10.20.5.0/26
ERROR: (NetcfgSubnetRangesOverlap) Subnet 'snet-x' is not valid in virtual network 'vnet-sysop'
because its IP address range overlaps with that of an existing subnet 'snet-pe'.

ERROR: (NetcfgSubnetRangeOutsideVnet) Subnet 'snet-x' is not valid because its IP address
range is outside the IP address range of virtual network 'vnet-sysop'.

A subnet must sit inside the VNet's address space and not overlap a sibling. And you cannot resize a subnet that has anything in it - a full AKS subnet means a new subnet and a migration, not an edit. Size generously up front.

Planning an address space

Peering is not transitive

Peering connects two VNets so their addresses can reach each other directly, as if they were one network. Transitive would mean "if A talks to B and B talks to C, then A talks to C". Peering is not:

hub  <-peer->  spoke-a
hub  <-peer->  spoke-b
spoke-a  X  spoke-b      (no route, unless traffic goes through a firewall/NVA in the hub)

Each peering is a direct link between two VNets. Spoke-to-spoke traffic needs a router in the hub (Azure Firewall, or an NVA - network virtual appliance, a VM running firewall or router software) plus route tables in the spokes pointing at it - lessons 23.7 and 23.14. (Hub and spoke: a central VNet everyone connects to, and the VNets hanging off it, like a bicycle wheel.)

The Terraform shape

resource "azurerm_subnet" "pe" {
  for_each             = { pe = "10.20.3.0/29", data = "10.20.5.0/26" }
  name                 = "snet-${each.key}"
  resource_group_name  = azurerm_resource_group.this.name
  virtual_network_name = azurerm_virtual_network.this.name
  address_prefixes     = [each.value]
}

The same subnets in Terraform (Ch 12): for_each makes one azurerm_subnet per map entry. for_each, not count - chapter 12's incident (12.20) was exactly this resource.

Reading a VNet end to end

$ az network vnet subnet show -g rg-oncall-lab --vnet-name vnet-sysop -n snet-app --query "{prefix:addressPrefix, nsg:networkSecurityGroup.id, rt:routeTable.id, pe:privateEndpointNetworkPolicies, delegations:delegations}"
{
  "delegations": [],
  "nsg": null,
  "pe": "Disabled",
  "prefix": "10.20.4.0/24",
  "rt": null
}

Everything a subnet can carry is on that object: the NSG (23.3), the route table (23.7), a delegation (a subnet handed to one service - App Service VNet integration, Azure Container Instances, PostgreSQL Flexible Server - that nothing else may use) and the private endpoint network policies flag (whether NSGs and route tables apply to private endpoints in it; enable it if you want to filter private endpoint traffic with NSGs).

Worked sizing examples

What each line says: the thing, how many addresses it needs, the prefix that fits. (Surge nodes are the extra nodes AKS adds temporarily during an upgrade, 23.28; 60 nodes x 31 is 31 addresses per node when every pod gets a subnet address, 23.22.)

private endpoints for ~20 services       20 + growth  -> /27 (27 usable)
App Gateway v2, autoscale to 125         Microsoft recommends a /24
Azure Firewall                           /26 minimum, name AzureFirewallSubnet
Bastion                                  /26 minimum, name AzureBastionSubnet
AKS overlay, 60 nodes + 10 surge         70 + 5 -> /25 (123 usable)
AKS node subnet, 60 nodes x 31           1,860 + surge -> /21

Round up a prefix, never down, and leave unallocated space in the VNet: the next subnet you will need is the one you did not plan for.

Common mistakes

$ az network vnet subnet create -g rg-oncall-lab --vnet-name vnet-sysop -n snet-db --address-prefixes 10.20.4.128/25
ERROR: (NetcfgSubnetRangesOverlap) Subnet 'snet-db' is not valid in virtual network 'vnet-sysop'
because its IP address range overlaps with that of an existing subnet 'snet-app'.

10.20.4.128/25 is the upper half of snet-app's 10.20.4.0/24. Write the ranges down in a table (or let Terraform's cidrsubnet() do it) before creating anything. A prefix that is not aligned to its size (10.20.4.64/25) is not a valid CIDR at all - the base of a /25 must be .0 or .128.

What you can now do

Why it helps

Address planning is the networking mistake you can't fix later. A VNet range that overlaps on-prem, another spoke or the AKS service CIDR means peerings that can't be created and routes that can't work, and the fix is rebuilding. A subnet sized too small for AKS means a new subnet, a new node pool and a migration during a busy week.

In practice you'll be asked to size subnets for private endpoints, App Gateway, Firewall and AKS, and to review Terraform that creates them. Knowing the five reserved addresses, the smallest allowed subnet, the required names like AzureFirewallSubnet, and to plan with cidrsubnet() and for_each rather than by hand, is what makes those reviews quick and right.

Commands in this lesson

az

FAQ

Why do I lose five addresses in every subnet?

Azure uses them: the first address is the network address, .1 is the default gateway, .2 and .3 map to Azure DNS, and the last address is broadcast. That's true whatever the prefix size. So a /24 has 251 usable addresses, a /28 has 11 and a /29, the smallest subnet Azure allows, has only 3. The first address you can assign is always .4.

Can I resize a subnet later?

Not while anything is in it. A subnet with NICs, private endpoints or AKS nodes can't be changed, which is why a full AKS subnet means creating a new subnet and a new node pool and migrating, not editing. You can add address space to the VNet itself, and in some cases resize empty subnets. So size generously at the start and leave unallocated space in the VNet for the subnet you didn't plan for.

Why do overlapping ranges matter so much?

Because two networks with overlapping addresses can't be peered or routed together: Azure refuses the peering with VnetAddressSpacesOverlap, and routing can't tell which side an address belongs to. That includes on-prem ranges, partner networks, other spokes and the AKS service and pod CIDRs, whose defaults like 10.0.0.0/16 overlap many corporate networks. Once workloads are running in overlapping ranges, the fix is re-addressing, which is very expensive.

What is a subnet delegation?

Handing a subnet to one Azure service, such as App Service VNet integration, Azure Container Instances or PostgreSQL Flexible Server, which then manages it and places its own resources there. Nothing else may be deployed into a delegated subnet. It shows up in delegations on the subnet object. Plan dedicated subnets for services that need delegation, just as for App Gateway and Firewall.

Why for_each instead of count for subnets in Terraform?

With count, each subnet is identified by its index. Remove one from the middle of the list and every subnet after it shifts index, so Terraform wants to destroy and recreate them, which fails or causes an outage for subnets in use. With for_each over a map, each subnet is identified by its key, like pe or data, and adding or removing one touches only that one. That was chapter 12's incident.

In an interview Mid

How would you plan the IP address space for a new Azure environment?

Also asked: How many usable IP addresses does an Azure subnet have, and why? · Is VNet peering transitive? · Why can you not simply make a subnet bigger later?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.