OnCallReady

Lesson 23.7 · Azure II: Networking & AKS · 13 min read

Route tables: where packets go after the NSG says yes

In plain words

Imagine a postal sorting office with a board of rules: "letters for this street go to van A, letters for this town go to van B, everything else goes to the airport". The most specific rule wins: a letter for your exact street takes van A even though it's also in the town. If two rules are equally specific, the one the manager added by hand beats the printed one.

A route table does that for packets leaving a subnet. Azure already has system routes: the VNet stays local, peered ranges go over peering, 0.0.0.0/0 goes to the internet, and some private ranges are dropped. A user-defined route like 0.0.0.0/0 -> VirtualAppliance 10.20.254.4 sends everything else to the firewall. Longest prefix wins, and on a tie, user beats BGP beats system.

The bank's rule: no workload talks to the internet directly; everything goes out through a firewall that allows only named destinations and logs every connection. An NSG can allow or deny, but it cannot send traffic somewhere. Routes do. This lesson is Azure's routing table, the same idea as ip route on your VM (8.11).

What you need to know already: 8.11 (routing tables, longest prefix match), 23.1 (VNets, subnets, peering, NVA), 23.3 (NSGs).

System routes

Every subnet has system routes Azure creates for you (destination range, then next hop - where the packet is sent next):

10.20.0.0/16      VnetLocal         (the VNet's own space)
<peered ranges>   VNetPeering
0.0.0.0/0         Internet
10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16, 100.64.0.0/10   None  (dropped unless something more specific exists)

None means drop. The private ranges (10/8, 172.16/12, 192.168/16, and 100.64/10, the "shared" range) are dropped unless your VNet, a peering or a route of yours covers them.

A route table is an Azure resource holding your own routes - each one a UDR (user-defined route). Attached to a subnet, it adds or overrides routes for traffic leaving that subnet.

Choosing a route

  1. Longest prefix match first: 10.20.4.0/24 beats 10.20.0.0/16 beats 0.0.0.0/0.
  2. On a tie in prefix length: user-defined > BGP > system. (BGP is the protocol routers use to tell each other which ranges they can reach; here, routes learnt from a VPN or ExpressRoute gateway, 23.1.)

So a UDR 0.0.0.0/0 -> VirtualAppliance 10.20.254.4 sends internet-bound traffic to a firewall, while traffic inside the VNet still matches the more specific 10.20.0.0/16 VnetLocal system route and goes direct.

Next hop types

VirtualAppliance        an IP: Azure Firewall's private IP or an NVA (needs --next-hop-ip-address)
VirtualNetworkGateway   the VPN / ExpressRoute gateway
VnetLocal               stay inside the VNet
Internet                out via Azure's internet edge
None                    drop (a black hole)

Three commands: az network route-table create makes an empty route table; az network route-table route create adds one route (--address-prefix the destination range, --next-hop-type, and --next-hop-ip-address for an appliance); az network vnet subnet update ... --route-table attaches it to a subnet.

az network route-table create -g rg-oncall-lab -n rt-egress
az network route-table route create -g rg-oncall-lab --route-table-name rt-egress -n default \
  --address-prefix 0.0.0.0/0 --next-hop-type VirtualAppliance --next-hop-ip-address 10.20.254.4
az network vnet subnet update -g rg-oncall-lab --vnet-name vnet-sysop -n snet-app --route-table rt-egress
# rt-egress = the route table of the next mission
az network route-table route list -g rg-oncall-lab --route-table-name rt-egress -o table
AddressPrefix  HasBgpOverride  Name     NextHopIpAddress  NextHopType       ProvisioningState  ResourceGroup
-------------  --------------  -------  ----------------  ----------------  -----------------  -------------
0.0.0.0/0      False           default  10.20.254.4       VirtualAppliance  Succeeded          rg-oncall-lab

Why you would do this

Egress control (control over traffic leaving your network). Banks do not let workloads reach the internet directly. All outbound traffic goes through Azure Firewall (or a third-party NVA) that allows only named destinations (*.docker.io no, mcr.microsoft.com yes, your package mirror yes) and logs everything.

For an AKS cluster (23.20) this is the create option --outbound-type userDefinedRouting: the cluster does not create an outbound public IP on its load balancer, and relies on your route table to send egress to the firewall. The firewall must then allow the endpoints AKS needs (the control plane FQDN, mcr.microsoft.com, *.data.mcr.microsoft.com, management.azure.com, login.microsoftonline.com, package repos for node images, NTP - the time service...). An FQDN (fully qualified domain name, Ch 8) is a complete host name like mcr.microsoft.com; firewall FQDN rules allow by name instead of by IP. Miss one and nodes fail to join or image pulls hang - one of the classic AKS-behind-a-firewall outages.

The mistakes

Seeing what really applies

For a VM NIC, az network nic show-effective-route-table prints the merged result of system, BGP and user routes with their source - the routing equivalent of listing NSG rules with the defaults. The drill does the longest-prefix reasoning by hand.

System routes you can see

# an illustration (no ▶): a subnet whose route table sends 0.0.0.0/0 to a firewall (abridged)
$ az network nic show-effective-route-table -g rg-oncall-lab -n vm-jump-nic -o table
Source    State    Address Prefix    Next Hop Type     Next Hop IP
--------  -------  ----------------  ----------------  -------------
Default   Active   10.20.0.0/16      VnetLocal
User      Active   0.0.0.0/0         VirtualAppliance  10.20.254.4
Default   Invalid  0.0.0.0/0         Internet
Default   Active   157.59.0.0/16     None
Default   Active   127.0.0.0/8       None

Run it on the lab's vm-jump-nic to see the real list for its subnet. Note the Invalid default route: your 0.0.0.0/0 replaced it. And the None routes: without a user default route, Azure drops packets to the private ranges (10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16, 100.64.0.0/10) that are not in your VNet or a peering - a packet to 10.99.0.1 matches the /8 None route, which is more specific than the /0 Internet route. Once a user 0.0.0.0/0 points at an appliance or a VPN gateway, Azure removes those four None routes, so 10.99.0.1 goes to the firewall (Microsoft Learn, "Azure virtual network traffic routing", routing example, Subnet1). On a tie of prefix length the user route wins.

Worked decisions

destination     matching routes                            winner
10.20.4.9       10.20.0.0/16 VnetLocal, 0/0 appliance      VnetLocal (/16)
8.8.8.8         0/0 appliance                              appliance
192.168.1.1     192.168.0.0/16 None, 192.168.0.0/16 gw     gateway (user wins tie)
10.99.0.1       0/0 appliance (the /8 None is removed)     appliance
10.99.0.1       no user 0/0: 10.0.0.0/8 None, 0/0 Internet None (/8 beats /0)

What you can now do

Why it helps

At a bank, workloads don't go to the internet directly: all egress goes through a firewall that allows only named destinations. That's a route table on every spoke subnet, and for AKS it's --outbound-type userDefinedRouting. When the firewall doesn't allow an endpoint AKS needs, nodes fail to join or image pulls hang, a classic outage you'll recognise.

Routing is also behind the hardest connectivity bugs: asymmetric paths that a stateful firewall silently drops, a private-range packet that goes nowhere because a system /8 beats your /0, or a route table attached to the firewall's own subnet. Reasoning through longest-prefix match by hand, as the drill trains, lets you find those in minutes instead of opening a ticket with Microsoft.

Commands in this lesson

az

FAQ

How does Azure pick a route when several match?

Longest prefix match first: a /24 beats a /16 beats a /8 beats a /0. Only on a tie in prefix length does the source matter: user-defined beats BGP, which beats system routes. So a UDR for 0.0.0.0/0 doesn't affect traffic inside the VNet, because the system 10.20.0.0/16 VnetLocal route is more specific. az network nic show-effective-route-table shows the merged result for a VM's NIC.

What is asymmetric routing and why does it break things?

Traffic takes one path out and a different path back, for example through the firewall on the way out but directly on the way back. A stateful firewall only sees half the conversation, considers the packets invalid and drops them, so connections hang. Typical causes are a UDR on one side only, like spoke A routing to the firewall but spoke B not, or a public IP on a VM that lets inbound traffic bypass the firewall.

Why is my packet to 10.99.0.1 dropped, and does a default route to the firewall change that?

Azure has system routes with next hop None for the private ranges 10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16 and 100.64.0.0/10: a private address that is not in your VNet or a peering matches the /8, which is more specific than the /0 Internet route, so the packet is dropped. A user route for 0.0.0.0/0 to an appliance or a VPN gateway makes Azure remove those None routes, and the packet then goes to the firewall. az network nic show-effective-route-table shows which case you are in.

What does outboundType userDefinedRouting do in AKS?

It tells AKS not to create an outbound public IP and SNAT rule on its load balancer, and to rely on your route table for egress instead, typically a 0.0.0.0/0 route to Azure Firewall. The subnet must have that route table before the cluster is created, and the firewall must allow everything AKS needs: the control plane FQDN, mcr.microsoft.com, management.azure.com, login.microsoftonline.com, node image repositories and NTP. Missing one breaks node provisioning or image pulls.

Which subnets should not get the default route to the firewall?

The firewall's own subnet, AzureFirewallSubnet, since routing it through itself loops. Also subnets whose platform services need direct internet access for their management traffic, such as Application Gateway v2 and some managed services like SQL Managed Instance, where a 0/0 to a firewall breaks provisioning or health. And watch --disable-bgp-route-propagation: it's needed for forced tunnelling, but it also stops VPN-learnt routes reaching the subnet.

In an interview Mid

How does Azure decide where a packet goes, and how do you send a subnet's internet traffic through a firewall?

Every subnet has system routes: the VNet's range VnetLocal, peered ranges VNetPeering, 0.0.0.0/0 Internet, and the private ranges (10/8, 172.16/12, 192.168/16, 100.64/10) to None - dropped - unless something more specific covers them.

Choosing: longest prefix match first; on a tie, user-defined > BGP > system.

To force egress through a firewall: a route table with a UDR 0.0.0.0/0 -> VirtualAppliance <firewall private IP>, attached to the subnet (az network route-table route create ..., az network vnet subnet update --route-table). Traffic inside the VNet still matches the more specific VnetLocal and goes direct. A user 0/0 to an appliance also removes the private-range None routes.

Mistakes to avoid: asymmetric routing (out via the firewall, back direct - the stateful firewall drops it), routing AzureFirewallSubnet through itself, and breaking subnets that need direct internet (App Gateway v2). For a cluster behind it, the firewall must allow the endpoints the nodes need, or they never join. az network nic show-effective-route-table shows what really applies.

Also asked: What is a user-defined route and why would you use one? · Cluster nodes behind a firewall fail to join after creation. What would you check? · What is asymmetric routing and why does a firewall drop it?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.