Chapter 23 Azure II: Networking & AKS
Subnets and the 5 reserved addresses, NSGs, routes, private endpoints and their DNS, load balancers, and AKS from the Azure side.
In plain words
Think of a gated town. Each neighbourhood has its own streets and house numbers; a few numbers on every street belong to the town and can't be used. Each neighbourhood has a guard at its entrance with a list of who may come in. Road signs decide which way cars go after the guard waves them through. Some shops have a private back door on your street so you never need the main road. Neighbourhoods connect only through the town centre, never directly to each other.
Azure networking is that town. VNets and subnets are the neighbourhoods and streets, NSGs are the guards, route tables are the road signs, private endpoints are the back doors, and hub and spoke is the town centre. AKS is a big neighbourhood of its own, and this chapter shows how it takes addresses, scales and lets people in.
Why it matters on call
AKS at a bank runs inside exactly this kind of network: a spoke VNet peered to a hub with a firewall, forced egress through a route table, private endpoints for Key Vault and storage, an Application Gateway in front. When something doesn't connect, the cause is almost always one of this chapter's pieces: a subnet out of IPs, an NSG rule with the wrong priority, a missing UDR on the return path, a private DNS zone not linked, a half-created peering.
You'll get these tickets weekly and review the Terraform for them. The AKS lessons cover the cluster decisions you can't undo later: CNI mode and subnet size, pool design, spot, and access through Entra ID with local accounts off. It comes after Azure identity because every one of these resources is created and accessed through it, and before the monitoring chapter, where you'll debug all of this from logs.
Lessons
- VNets and subnets: the five addresses you never get
- Network security groups: priorities, defaults, service tags
- Route tables: where packets go after the NSG says yes
- Private endpoints and the DNS that makes them work
- Hub and spoke: peering, transit, and why spokes cannot see each other
- Load Balancer vs Application Gateway, and the idle timeout
- AKS from the Azure side
- AKS networking: node subnet, overlay, and the IP bill
- Cluster autoscaler, spot pools, and pool design
- AKS access: Entra ID, Azure RBAC for Kubernetes, local accounts
22 hands-on labs (missions, incidents and drills) run in the terminal: Open this chapter in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.
Questions people ask
Is Azure networking similar to AWS or on-prem networking?
The concepts carry over: private address spaces, subnets, stateful filters, route tables, load balancers, peering. The details differ. Azure reserves five addresses per subnet, NSGs have default rules that allow all traffic inside the VNet, peering is never transitive, and private access to PaaS services depends on private DNS zones. Much of what goes wrong for people coming from on-prem is assuming a router in the middle that Azure doesn't have unless you add one.
Why is DNS so central in this chapter?
Because private endpoints only work if the service's normal name resolves to the private IP from inside your network. The application still connects to stsysoplab.blob.core.windows.net; a private DNS zone linked to the VNet makes that name return 10.20.3.4. If the zone isn't linked, or a custom DNS server doesn't forward, the name resolves to the public IP, and with public access disabled the connection fails with an error that looks like an auth problem.
Do I need to know this if AKS manages its own network?
Yes. AKS manages some things, the load balancer, its NSG and node VMs in the MC_ resource group, but it lives in your subnet, uses your route tables and DNS, and is limited by your IP plan. Subnet sizing, egress through a firewall, private endpoints for its dependencies and access through Entra ID are your decisions, and several of them can't be changed after the cluster is created.
Why hub and spoke instead of connecting everything directly?
Centralisation and control. The hub holds shared services, the firewall, VPN or ExpressRoute gateways, DNS resolvers, Bastion, so every spoke doesn't need its own. Traffic between spokes and to the internet goes through the firewall, where it's filtered and logged, which is what security and audit teams require. A full mesh of peerings would bypass that inspection and grow quadratically as spokes are added.
What does the lab simulate versus real Azure?
The CLI commands, output shapes and error codes are the real ones, and the resources behave like the real ones for what the lessons teach: subnet capacity, NSG evaluation, route selection, private DNS resolution, peering states, AKS pool and access errors. There is no real traffic across Azure's network, and some things, like a VM's effective routes, are illustrative. Anything simulator-only is labelled (simulator).