Azure II: Networking & AKS: interview questions
The question you are most likely to get for each topic, a model answer, and what else comes up. From chapter 23 of the course.
Explain hub-and-spoke network architecture in Azure. Mid
One central hub VNet holds the shared services - Azure Firewall (or an NVA), the VPN/ExpressRoute gateway to on-prem, DNS resolvers, Bastion. Each workload or environment gets its own spoke VNet, peered only with the hub.
Why: one place to control and log egress, one link to on-prem, one DNS setup, and teams isolated from each other by default.
The rules that make it work:
- Peering is not transitive. spoke-a <-> hub and hub <-> spoke-b does not connect the spokes. Spoke-to-spoke goes through the hub firewall: a UDR in each spoke pointing at the firewall's IP (
VirtualAppliance), the firewall allowing it, andallowForwardedTrafficon the peerings. - A peering is two objects; half of one is
Initiatedand carries nothing. - Address spaces must not overlap anywhere - plan them centrally.
--allow-gateway-transit/--use-remote-gatewayslet spokes use the hub's gateway.- DNS: one central private DNS zone per service, linked to every spoke, or private endpoints resolve to public IPs.
Also asked: A pod cannot reach a database behind a private endpoint in another spoke. How do you troubleshoot it? · What are the main building blocks of an Azure virtual network? · How would you size the subnet for an AKS cluster on Azure CNI?
How would you plan the IP address space for a new Azure environment? Mid
- Never overlap anything you may ever peer with or route to: other VNets, on-prem, partner networks, and the cluster's service CIDR (default 10.0.0.0/16!) and pod CIDR. Overlaps are the one mistake you cannot fix later - peering refuses them.
- One subnet per purpose and security boundary: cluster nodes, App Gateway (its own, /24 recommended), private endpoints, VMs, databases. Special subnets have fixed names and minimum sizes:
AzureFirewallSubnet/26,AzureBastionSubnet/26,GatewaySubnet. - Do the Azure arithmetic: each subnet loses 5 addresses (.0 network, .1 gateway, .2-.3 Azure DNS, the last broadcast). Usable =
2^(32-n) - 5; the first address you get is .4; a /29 holds three things. - Size generously and round up: you cannot resize a subnet that has anything in it - a full subnet means a new subnet and a migration. Leave unallocated space in the VNet for the subnet you did not plan for.
- Write the plan down (or let Terraform's
cidrsubnet()compute it) before creating anything.
Also asked: How many usable IP addresses does an Azure subnet have, and why? · Is VNet peering transitive? · Why can you not simply make a subnet bigger later?
Learn it: 23.1 VNets and subnets: the five addresses you never get
What is a network security group and how are its rules evaluated? Mid
An NSG is a list of allow/deny rules on source, source port, destination, destination port and protocol, attached to a subnet or a NIC. It works at L4 (addresses and ports, not URLs) and is stateful - replies to an allowed connection are allowed automatically.
Evaluation, per direction: lowest priority number first, first match wins. Your rules use 100-4096; below them sit default rules you cannot delete:
AllowVnetInBound- everything in the VNet (including peered VNets and on-prem) reaches everything. An empty NSG is "deny the internet", not "deny all".AllowAzureLoadBalancerInBound- health probes from168.63.129.16.- Outbound to the internet allowed;
DenyAllInBoundat 65500.
So isolating a subnet = specific allows at low numbers, then your own "deny VirtualNetwork" above the defaults (e.g. 4000). Use service tags (AzureLoadBalancer, Storage) instead of IP lists. Gotcha: az network nsg rule create without --destination-port-ranges defaults to 80. With NSGs on both subnet and NIC, both must allow.
Also asked: How would you isolate an application subnet so only the App Gateway subnet can reach it on 443? · What are service tags and why use them in NSG rules? · What can an NSG not filter, and what do you use instead?
Learn it: 23.3 Network security groups: priorities, defaults, service tags
How does Azure decide where a packet goes, and how do you send a subnet's internet traffic through a firewall? Mid
Every subnet has system routes: the VNet's range VnetLocal, peered ranges VNetPeering, 0.0.0.0/0 Internet, and the private ranges (10/8, 172.16/12, 192.168/16, 100.64/10) to None - dropped - unless something more specific covers them.
Choosing: longest prefix match first; on a tie, user-defined > BGP > system.
To force egress through a firewall: a route table with a UDR 0.0.0.0/0 -> VirtualAppliance <firewall private IP>, attached to the subnet (az network route-table route create ..., az network vnet subnet update --route-table). Traffic inside the VNet still matches the more specific VnetLocal and goes direct. A user 0/0 to an appliance also removes the private-range None routes.
Mistakes to avoid: asymmetric routing (out via the firewall, back direct - the stateful firewall drops it), routing AzureFirewallSubnet through itself, and breaking subnets that need direct internet (App Gateway v2). For a cluster behind it, the firewall must allow the endpoints the nodes need, or they never join. az network nic show-effective-route-table shows what really applies.
Also asked: What is a user-defined route and why would you use one? · Cluster nodes behind a firewall fail to join after creation. What would you check? · What is asymmetric routing and why does a firewall drop it?
Learn it: 23.7 Route tables: where packets go after the NSG says yes
After disabling public network access on a storage account or Key Vault, apps inside the VNet get 403. How do you diagnose it? Mid
Almost always DNS. A private endpoint gives the service a private IP in your subnet, but the app still uses the public name (stsysoplab.blob.core.windows.net), which must resolve to that IP from inside the VNet. Azure does it with a CNAME to *.privatelink.blob.core.windows.net, answered privately only where a matching private DNS zone is visible. Otherwise the name resolves to the public IP - which is now closed, hence a 403 that looks like auth.
- From inside the VNet:
nslookup <name>- a public IP means the chain did not stop at the private zone. az network private-dns link vnet list ...- is the zone linked to this VNet?az network private-dns record-set a list ...- does the A record exist, and does it match the endpoint's IP? (A zone group keeps it in sync; hand-made records go stale.)- Custom DNS servers must forward
privatelink.*to Azure DNS168.63.129.16. - One endpoint per sub-resource: a
blobendpoint does nothing forfile.
Also asked: What is an Azure private endpoint? · What is the difference between a private endpoint and a service endpoint? · How would you organise private DNS zones across many subscriptions?
Learn it: 23.10 Private endpoints and the DNS that makes them work
Two spokes cannot communicate through the hub firewall. How do you troubleshoot it? Mid
Peering is not transitive, so spoke-to-spoke needs four pieces, and I check them in order:
az network vnet peering liston both spokes and the hub - every peeringConnected?Initiated= only one side exists, nothing flows.allowForwardedTrafficon the destination spoke's peering - traffic arriving from the firewall did not originate in the hub.- Route table on the source subnet: destination spoke's prefix ->
VirtualAppliancethe firewall's IP. Without it the packet never goes to the hub. - Route table on the destination subnet for the way back. Without it replies go nowhere, or around the firewall, which drops the asymmetric half-flow.
- The firewall rules allow the flow (its logs show allowed/denied).
- The NSGs on both subnets.
If the network is fine but a name is not, it is DNS: the private DNS zone must be linked to (or resolvable from) the calling spoke.
Also asked: Is VNet peering transitive, and what does that mean in practice? · What does gateway transit do in a hub-and-spoke network? · Why does a private endpoint in one spoke often not resolve from another spoke?
Learn it: 23.14 Hub and spoke: peering, transit, and why spokes cannot see each other
What is the difference between Azure Load Balancer and Application Gateway, and how does each typically fail? Mid
- Azure Load Balancer - L4: forwards TCP/UDP connections by IP and port (a 5-tuple hash). Passes TLS through. Every
type: LoadBalancerService on a cluster uses one. - Application Gateway - L7 managed reverse proxy: reads host, path and headers, terminates TLS, rewrites and redirects, and can run a WAF.
Both send traffic only to backends whose health probe passes.
Typical failures:
- App Gateway 502 - no healthy backend: probe path, host header, port or expected status do not match what the backend answers. Read
az network application-gateway show-backend-healthfirst. 504 = backend too slow; 403 = WAF blocked it. The Server header says whether the gateway itself answered. - Load Balancer idle timeout - a flow with no packets for 4 minutes (default) is forgotten. Without TCP reset the next packet is silently dropped and a pooled connection hangs. Fix on the client: TCP keepalives or pool idle eviction below the timeout.
- SNAT exhaustion - many short outbound connections exhaust ports; fix with connection pooling or a NAT Gateway.
Also asked: Long-lived connections from pods hang after quiet periods. What is happening and how do you fix it? · Users get intermittent 502 errors from an Application Gateway. How do you investigate? · What is SNAT port exhaustion and how do you avoid it?
Learn it: 23.16 Load Balancer vs Application Gateway, and the idle timeout
What does Azure manage for you in AKS, and what do you manage? Mid
Azure runs the control plane: API server, etcd, scheduler, controller manager - in Microsoft's subscription. You see one endpoint (fqdn), never the machines; no etcd backups or certificate rotation of your own. The tier decides the SLA (Free: none; Standard: the production default).
You own the nodes: each node pool is a VM scale set you pay for, in the node resource group (MC_<rg>_<cluster>_<region>) together with the load balancer, public IPs, PVC disks and the kubelet identity. That group is AKS-managed - change it through az aks or Terraform, never by hand.
Also yours:
- Pool design - a System pool (tainted
CriticalAddonsOnly, 2-3 nodes across zones) and User pools per workload shape. VM size and maxPods are fixed at creation: change = new pool, drain, delete. - Upgrades - Kubernetes minor by minor (control plane, then pools) plus node images; surge nodes need free IPs, and a PDB with
maxUnavailable: 0stalls drains. Auto-upgrade channels plus a maintenance window. - The network, identities, and the bill.
Also asked: How would you design node pools for a production AKS cluster? · How do you run AKS upgrades safely in production? · Why should you not change resources in the MC_ resource group by hand?
Learn it: 23.20 AKS from the Azure side
How would you size a subnet for an AKS cluster using Azure CNI? Mid
First ask which Azure CNI:
- Overlay (the default) - pods get IPs from a private pod CIDR (/24 per node), so the node subnet needs one IP per node. Sizing is easy; this is what almost everyone wants.
- Node subnet (legacy) - each node pre-allocates its own IP plus maxPods pod IPs at creation, used or not:
IPs = (nodes + surge) x (maxPods + 1) + 5 reserved
50 nodes, maxPods 30, surge 1: 51 x 31 = 1,581 + 5 = 1,586
/22 = 1,019 usable too small /21 = 2,043 usable fits
Count every pool in the subnet at its max (autoscaler maxCount), plus surge for upgrades - a subnet filled to the last node makes the next upgrade fail. Too small fails up front with InsufficientSubnetSize, and the message is the formula.
You cannot resize a used subnet or change a pool's subnet or maxPods, so the way out is a new subnet, a new pool, drain the old one - or Overlay on a new cluster. Also: the service CIDR (default 10.0.0.0/16) must not overlap your networks.
Also asked: Compare Azure CNI Overlay and Azure CNI with a node subnet. When would you choose each? · What happens when an AKS cluster runs out of IP addresses in its subnet? · Why must the cluster's service CIDR not overlap your corporate networks?
Learn it: 23.22 AKS networking: node subnet, overlay, and the IP bill
Pods are Pending but the cluster autoscaler isn't adding nodes. What do you check? Mid
The cluster autoscaler adds nodes only for pods that are Pending because nothing fits, judged by requests, not usage. It writes its reason as events on the pod: kubectl describe pod / kubectl get events.
NotTriggerScaleUp ... max node group size reached- the pool is at--max-count; raise it.NotTriggerScaleUp ... untolerated taint,Insufficient memory- no pool could ever fit it: a taint (e.g. the spot taint) without toleration, a node selector or affinity, or a pod bigger than the VM size.FailedScaleUp ... InsufficientSubnetSize- Azure refused: the subnet is full (node-subnet CNI), the vCPU quota is used up, or no spot capacity.- Is the autoscaler enabled on that pool (it is per pool,
az aks nodepool update --enable-cluster-autoscaler)? - Pods without requests confuse it completely.
And timing: a new node takes minutes to boot, so short spikes are the HPA's and headroom's job. Scale-down is blocked by PDBs, emptyDir pods, bare pods and safe-to-evict: "false".
Also asked: How does the cluster autoscaler decide to add or remove nodes? · How would you use spot node pools without risking availability? · What is the difference between the cluster autoscaler and the Horizontal Pod Autoscaler?
Learn it: 23.25 Cluster autoscaler, spot pools, and pool design
How do users get access to an Entra-integrated AKS cluster, and what does a developer need to deploy to one namespace? Mid
Two separate gates:
- ARM - may you get a kubeconfig?
az aks get-credentialsneedslistClusterUserCredential, from Azure Kubernetes Service Cluster User Role. The kubeconfig holds no credential, only an exec plugin: kubelogin fetches an Entra token each time (kubelogin convert-kubeconfig -l azureclireuses youraz login). Missing it:AuthorizationFailed. - Kubernetes API - may your token do this? Either Kubernetes RBAC with Entra group object ids as subjects, or Azure RBAC for Kubernetes (
--enable-azure-rbac): role assignments like RBAC Reader/Writer/Admin, scoped to the cluster or to.../namespaces/shop. They are dataActions, so even a subscription Owner getsForbidden("User does not have access to the resource in Azure").
So a shop developer needs Cluster User Role + RBAC Writer on namespaces/shop, both granted to the team's group.
Then close the back door: --admin credentials bypass Entra and RBAC, so --disable-local-accounts - after setting up a break-glass group with Cluster Admin via PIM.
Also asked: Compare Kubernetes RBAC and Azure RBAC for Kubernetes authorization in AKS. · Why would you disable local accounts on an AKS cluster, and what must exist first? · How should automation authenticate to an AKS cluster?
Learn it: 23.31 AKS access: Entra ID, Azure RBAC for Kubernetes, local accounts
Practise these answers with flashcards and labs Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.