OnCallReady

Lesson 23.20 · Azure II: Networking & AKS · 12 min read

AKS from the Azure side

In plain words

Imagine renting a flat in a managed building. The landlord runs the boiler, the lifts and the security desk; you never see the machinery and you can't go into the boiler room. You furnish your own flat, pay for its electricity, and decide how many rooms you rent. There's also a storage room in the basement labelled with your name, which the landlord manages for you: if you rearrange things in there yourself, the landlord's staff put them back, or something breaks.

AKS is that building. Microsoft runs the control plane, the API server and etcd, and you only see its endpoint. Your node pools are VM scale sets you pay for. The basement room is the node resource group, MC_rg-oncall-lab_aks-sysop_westeurope, which holds the VMs, load balancer, disks and kubelet identity, and is managed by AKS, not by you.

Running Kubernetes yourself means running its control plane: etcd backups, certificate rotation, kubeadm upgrades (Ch 18). AKS (Azure Kubernetes Service) is Azure running that part for you, while you still own the nodes, their network and their bill. Chapters 15-19 taught Kubernetes; this lesson is what Azure does around it.

What you need to know already: 15.5 (the control plane: API server, etcd, scheduler, controller manager), 15.7 (nodes, kubelet), 18.15-18.16 (version skew, upgrades), 18.25 (cordon and drain), 17.26 (PDBs), 23.1 (subnets), 22.1 (resource groups), 22.9 (managed identities).

What you get

az aks show -g <rg> -n <cluster> prints the whole cluster object; the --query picks the version, pricing tier, API address, node resource group and whether it is running:

$ az aks show -g rg-oncall-lab -n aks-sysop --query "{version:kubernetesVersion, tier:sku.tier, fqdn:fqdn, nodeRg:nodeResourceGroup, power:powerState.code}"
{
  "fqdn": "aks-sysop-dns-87ec924b.hcp.westeurope.azmk8s.io",
  "nodeRg": "MC_rg-oncall-lab_aks-sysop_westeurope",
  "power": "Running",
  "tier": "Standard",
  "version": "1.33.3"
}

The MC_ group is AKS-managed. Changing things in it by hand (scaling the VMSS, editing the LB, deleting a disk the cluster still references) puts the cluster out of sync with what AKS thinks it has; the next reconcile (15.9) undoes it or breaks. Operate through az aks / Terraform.

A VM's size (its SKU, such as Standard_D8ds_v5) fixes its CPUs and memory: in that name, D = general purpose family, 8 = vCPUs, ds = local SSD and premium storage, v5 = generation.

Node pools

$ az aks nodepool list -g rg-oncall-lab --cluster-name aks-sysop -o table
Name    OsType  KubernetesVersion  VmSize            Count  MaxPods  ProvisioningState  Mode
------  ------  -----------------  ----------------  -----  -------  -----------------  ------
system  Linux   1.33.3             Standard_D4ds_v5  2      30       Succeeded          System
user    Linux   1.33.3             Standard_D8ds_v5  3      30       Succeeded          User

az aks nodepool list -g <rg> --cluster-name <cluster> -o table lists the pools. Mode is the column to read:

Zones

An availability zone is a physically separate data centre inside one region (own power and cooling); a region usually has three. --zones 1 2 3 spreads a pool's VMs across them. Remember from the storage lesson that Azure Disks are zonal: a StatefulSet pod whose disk is in zone 1 can only run on a zone-1 node.

Upgrades

Two layers, both yours to schedule:

  1. Kubernetes version: az aks upgrade - control plane first, then each pool. AKS supports roughly the three newest minor versions (the middle number: 1.33.3); fall behind and you are forced up. Minor versions must go one at a time (18.15).
  2. Node image: the OS image of the nodes (security patches) - az aks nodepool upgrade --node-image-only, or an auto-upgrade channel.

Pools upgrade by surge: AKS adds a new node, cordons and drains an old one (respecting PodDisruptionBudgets), deletes it, repeats. maxSurge (default 10% on new pools, at least one node) controls how many extra nodes at once. Surge nodes need IP addresses - on Azure CNI with a full subnet, an upgrade fails before it starts. A PDB with maxUnavailable: 0 (or minAvailable equal to replicas) makes drains hang forever and the upgrade stalls.

Auto-upgrade channels (patch, stable, node-image - settings that let AKS upgrade on its own) plus a planned maintenance window (the hours AKS is allowed to do it) is the production pattern: patches arrive on their own, in a window you chose.

Terraform shape

The same cluster in Terraform (Ch 12-14). Some settings in it are taught later in this chapter: azure_active_directory_role_based_access_control and local_account_disabled in 23.31, network_profile in 23.22; the OIDC and workload identity lines are 22.12's.

resource "azurerm_kubernetes_cluster" "this" {
  name                = "aks-sysop"
  resource_group_name = azurerm_resource_group.this.name
  location            = "westeurope"
  dns_prefix          = "aks-sysop"
  sku_tier            = "Standard"
  oidc_issuer_enabled       = true
  workload_identity_enabled = true
  azure_active_directory_role_based_access_control { azure_rbac_enabled = true }
  local_account_disabled = true

  default_node_pool {
    name                         = "system"
    vm_size                      = "Standard_D4ds_v5"
    node_count                   = 2
    vnet_subnet_id               = azurerm_subnet.aks.id
    only_critical_addons_enabled = true
    zones                        = ["1", "2", "3"]
  }
  identity { type = "SystemAssigned" }
  network_profile { network_plugin = "azure", network_plugin_mode = "overlay" }
}

resource "azurerm_kubernetes_cluster_node_pool" "user" {
  name                  = "user"
  kubernetes_cluster_id = azurerm_kubernetes_cluster.this.id
  vm_size               = "Standard_D8ds_v5"
  auto_scaling_enabled  = true
  min_count             = 2
  max_count             = 6
}

Reading the MC_ group

$ az resource list -g MC_rg-oncall-lab_aks-sysop_westeurope --query "[].{Name:name, Type:type}" -o table
Name                                      Type
----------------------------------------  --------------------------------
kubernetes                                Microsoft.Network/loadBalancers

(abridged - a real one also holds aks-system-*-vmss and aks-user-*-vmss scale sets, the kubelet identity aks-sysop-agentpool, a public IP for outbound traffic, an NSG, a route table on kubenet (23.22), and one managed disk per PVC). Everything there is billed to you and deleted when the cluster is.

Two useful knobs on that group: --node-resource-group at create time gives it a name that is not MC_..., and node resource group lockdown (--nrg-lockdown-restriction-level ReadOnly) puts a deny assignment on it (an RBAC rule that blocks actions even for Owners, 22.13) so nobody can change AKS-managed resources by hand.

Stop, start, delete

az aks stop deallocates the nodes and the control plane (you stop paying for VMs; disks and IPs stay). az aks start brings it back - with new node VMs, so anything that assumed node IPs or local disk state is wrong afterwards. Great for dev clusters at night; never a way to "restart" production.

What you can now do

Why it helps

Knowing what Azure does around Kubernetes prevents a class of self-inflicted incidents: someone scales a VMSS by hand in the MC_ group, deletes a disk the cluster still references, or edits the load balancer, and AKS either reverts it or the cluster ends up broken. You'll know to operate through az aks and Terraform, and why node resource group lockdown exists.

It also covers the decisions you'll make or review: Standard tier for production SLAs, system pools tainted CriticalAddonsOnly, pool VM size and maxPods fixed at creation, zones, and upgrades. Upgrades fail in two classic ways, no IPs for surge nodes and PDBs that block drains, and both are CKA-adjacent interview favourites. az aks stop is great for dev at night and never a way to restart production.

Commands in this lesson

az

FAQ

Can I access the AKS control plane?

Only through its API endpoint, the cluster's FQDN. The API server, etcd, scheduler and controller manager run in Microsoft's subscription; you can't SSH to them, see their VMs or read etcd directly, and kubectl get nodes shows only your nodes. You get control plane logs through diagnostic settings, and the Standard tier adds an uptime SLA and more API server capacity. The Free tier has no SLA and is meant for development.

What is the MC_ resource group?

The node resource group AKS creates for the cluster's infrastructure: VM scale sets for each node pool, the load balancer, public IPs, the NSG and route table where applicable, managed disks for PVCs, and the kubelet identity. It's managed by AKS: changing things there by hand puts the cluster out of sync, and the next reconcile reverts or breaks it. It's deleted with the cluster. Node resource group lockdown can prevent manual changes.

Why can't I change a node pool's VM size?

VM size, OS disk settings, maxPods and the subnet are fixed when a pool is created, because they're baked into the scale set and the node's network configuration. To change them, create a new pool with the new settings, cordon and drain the old pool so workloads move, and delete it. Pools are cheap and designed to be replaced, so plan for it, and keep workloads tolerant of moving between nodes.

What is the difference between a Kubernetes upgrade and a node image upgrade?

A Kubernetes upgrade moves the cluster to a new Kubernetes version, control plane first and then each node pool, one minor version at a time. A node image upgrade replaces the nodes' OS image with the latest patched one without changing the Kubernetes version, for security fixes. Both roll through the pool with surge nodes and drains. Auto-upgrade channels and a planned maintenance window let both happen automatically at a time you choose.

Why did my AKS upgrade get stuck?

The two most common reasons: no IP addresses for surge nodes, because on node-subnet CNI a full subnet can't fit even one extra node with its maxPods allocation, so the upgrade fails before it starts; and PodDisruptionBudgets that don't allow any disruption, like maxUnavailable: 0 or minAvailable equal to replicas, which make the drain of each node wait forever. Check az aks show for the error and kubectl get pdb -A for PDBs with zero allowed disruptions.

In an interview Mid

What does Azure manage for you in AKS, and what do you manage?

Azure runs the control plane: API server, etcd, scheduler, controller manager - in Microsoft's subscription. You see one endpoint (fqdn), never the machines; no etcd backups or certificate rotation of your own. The tier decides the SLA (Free: none; Standard: the production default).

You own the nodes: each node pool is a VM scale set you pay for, in the node resource group (MC_<rg>_<cluster>_<region>) together with the load balancer, public IPs, PVC disks and the kubelet identity. That group is AKS-managed - change it through az aks or Terraform, never by hand.

Also yours:

Also asked: How would you design node pools for a production AKS cluster? · How do you run AKS upgrades safely in production? · Why should you not change resources in the MC_ resource group by hand?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.