Running Kubernetes yourself means running its control plane: etcd backups, certificate rotation, kubeadm upgrades (Ch 18). AKS (Azure Kubernetes Service) is Azure running that part for you, while you still own the nodes, their network and their bill. Chapters 15-19 taught Kubernetes; this lesson is what Azure does around it.
What you need to know already: 15.5 (the control plane: API server, etcd, scheduler, controller manager), 15.7 (nodes, kubelet), 18.15-18.16 (version skew, upgrades), 18.25 (cordon and drain), 17.26 (PDBs), 23.1 (subnets), 22.1 (resource groups), 22.9 (managed identities).
What you get
az aks show -g <rg> -n <cluster> prints the whole cluster object; the --query picks the version, pricing tier, API address, node resource group and whether it is running:
$ az aks show -g rg-oncall-lab -n aks-sysop --query "{version:kubernetesVersion, tier:sku.tier, fqdn:fqdn, nodeRg:nodeResourceGroup, power:powerState.code}"
{
"fqdn": "aks-sysop-dns-87ec924b.hcp.westeurope.azmk8s.io",
"nodeRg": "MC_rg-oncall-lab_aks-sysop_westeurope",
"power": "Running",
"tier": "Standard",
"version": "1.33.3"
}
- The control plane is Microsoft's. API server, etcd, scheduler and controller manager run in Azure's subscription. You see one endpoint (the
fqdn), you never see the machines, and you cannot SSH to them.kubectl get nodesshows only your nodes. - Tier: Free (no SLA, for dev), Standard (financially backed uptime SLA · an SLA (0.8) where Microsoft refunds money if it misses - and more API server capacity; the production default), Premium (adds long-term support versions).
- Nodes are VM scale sets (VMSS, 23.3: a group of identical VMs Azure grows and shrinks as one) you pay for, one scale set per node pool (a group of nodes with the same VM size and settings), living in the node resource group (
MC_<rg>_<cluster>_<region>) together with the load balancer, public IPs, managed disks for PVCs (22.31) and the kubelet identity (the managed identity the nodes use, for example to pull images).
The MC_ group is AKS-managed. Changing things in it by hand (scaling the VMSS, editing the LB, deleting a disk the cluster still references) puts the cluster out of sync with what AKS thinks it has; the next reconcile (15.9) undoes it or breaks. Operate through az aks / Terraform.
A VM's size (its SKU, such as Standard_D8ds_v5) fixes its CPUs and memory: in that name, D = general purpose family, 8 = vCPUs, ds = local SSD and premium storage, v5 = generation.
Node pools
$ az aks nodepool list -g rg-oncall-lab --cluster-name aks-sysop -o table
Name OsType KubernetesVersion VmSize Count MaxPods ProvisioningState Mode
------ ------ ----------------- ---------------- ----- ------- ----------------- ------
system Linux 1.33.3 Standard_D4ds_v5 2 30 Succeeded System
user Linux 1.33.3 Standard_D8ds_v5 3 30 Succeeded User
az aks nodepool list -g <rg> --cluster-name <cluster> -o table lists the pools. Mode is the column to read:
- System pools run CoreDNS (16.13), metrics-server, konnectivity (the tunnel between Azure's control plane and your nodes) and the other kube-system pods. At least one is required; give it
--node-taints CriticalAddonsOnly=true:NoScheduleso your apps stay off it, and at least 2-3 nodes across zones. - User pools run your workloads. Separate pools for separate shapes: general purpose, memory-heavy, GPU, spot, Windows.
- A pool's VM size, OS disk and maxPods are fixed at creation. Changing them means a new pool, cordon/drain the old one, delete it. Pools are cheap; plan to replace them.
Zones
An availability zone is a physically separate data centre inside one region (own power and cooling); a region usually has three. --zones 1 2 3 spreads a pool's VMs across them. Remember from the storage lesson that Azure Disks are zonal: a StatefulSet pod whose disk is in zone 1 can only run on a zone-1 node.
Upgrades
Two layers, both yours to schedule:
- Kubernetes version:
az aks upgrade- control plane first, then each pool. AKS supports roughly the three newest minor versions (the middle number: 1.33.3); fall behind and you are forced up. Minor versions must go one at a time (18.15). - Node image: the OS image of the nodes (security patches) -
az aks nodepool upgrade --node-image-only, or an auto-upgrade channel.
Pools upgrade by surge: AKS adds a new node, cordons and drains an old one (respecting PodDisruptionBudgets), deletes it, repeats. maxSurge (default 10% on new pools, at least one node) controls how many extra nodes at once. Surge nodes need IP addresses - on Azure CNI with a full subnet, an upgrade fails before it starts. A PDB with maxUnavailable: 0 (or minAvailable equal to replicas) makes drains hang forever and the upgrade stalls.
Auto-upgrade channels (patch, stable, node-image - settings that let AKS upgrade on its own) plus a planned maintenance window (the hours AKS is allowed to do it) is the production pattern: patches arrive on their own, in a window you chose.
Terraform shape
The same cluster in Terraform (Ch 12-14). Some settings in it are taught later in this chapter: azure_active_directory_role_based_access_control and local_account_disabled in 23.31, network_profile in 23.22; the OIDC and workload identity lines are 22.12's.
resource "azurerm_kubernetes_cluster" "this" {
name = "aks-sysop"
resource_group_name = azurerm_resource_group.this.name
location = "westeurope"
dns_prefix = "aks-sysop"
sku_tier = "Standard"
oidc_issuer_enabled = true
workload_identity_enabled = true
azure_active_directory_role_based_access_control { azure_rbac_enabled = true }
local_account_disabled = true
default_node_pool {
name = "system"
vm_size = "Standard_D4ds_v5"
node_count = 2
vnet_subnet_id = azurerm_subnet.aks.id
only_critical_addons_enabled = true
zones = ["1", "2", "3"]
}
identity { type = "SystemAssigned" }
network_profile { network_plugin = "azure", network_plugin_mode = "overlay" }
}
resource "azurerm_kubernetes_cluster_node_pool" "user" {
name = "user"
kubernetes_cluster_id = azurerm_kubernetes_cluster.this.id
vm_size = "Standard_D8ds_v5"
auto_scaling_enabled = true
min_count = 2
max_count = 6
}
Reading the MC_ group
$ az resource list -g MC_rg-oncall-lab_aks-sysop_westeurope --query "[].{Name:name, Type:type}" -o table
Name Type
---------------------------------------- --------------------------------
kubernetes Microsoft.Network/loadBalancers
(abridged - a real one also holds aks-system-*-vmss and aks-user-*-vmss scale sets, the kubelet identity aks-sysop-agentpool, a public IP for outbound traffic, an NSG, a route table on kubenet (23.22), and one managed disk per PVC). Everything there is billed to you and deleted when the cluster is.
Two useful knobs on that group: --node-resource-group at create time gives it a name that is not MC_..., and node resource group lockdown (--nrg-lockdown-restriction-level ReadOnly) puts a deny assignment on it (an RBAC rule that blocks actions even for Owners, 22.13) so nobody can change AKS-managed resources by hand.
Stop, start, delete
az aks stop deallocates the nodes and the control plane (you stop paying for VMs; disks and IPs stay). az aks start brings it back - with new node VMs, so anything that assumed node IPs or local disk state is wrong afterwards. Great for dev clusters at night; never a way to "restart" production.
What you can now do
- Say what Azure runs (the control plane) and what you pay for and operate (node pools as VM scale sets in the MC_ group).
- Read a cluster and its pools with
az aks showandaz aks nodepool list. - Plan an upgrade: minor by minor, control plane first, surge needs IPs, PDBs can stall drains.