OnCallReady

Azure II: Networking & AKS: interview questions

The question you are most likely to get for each topic, a model answer, and what else comes up. From chapter 23 of the course.

Explain hub-and-spoke network architecture in Azure. Mid

One central hub VNet holds the shared services - Azure Firewall (or an NVA), the VPN/ExpressRoute gateway to on-prem, DNS resolvers, Bastion. Each workload or environment gets its own spoke VNet, peered only with the hub.

Why: one place to control and log egress, one link to on-prem, one DNS setup, and teams isolated from each other by default.

The rules that make it work:

Also asked: A pod cannot reach a database behind a private endpoint in another spoke. How do you troubleshoot it? · What are the main building blocks of an Azure virtual network? · How would you size the subnet for an AKS cluster on Azure CNI?

How would you plan the IP address space for a new Azure environment? Mid

Also asked: How many usable IP addresses does an Azure subnet have, and why? · Is VNet peering transitive? · Why can you not simply make a subnet bigger later?

Learn it: 23.1 VNets and subnets: the five addresses you never get

What is a network security group and how are its rules evaluated? Mid

An NSG is a list of allow/deny rules on source, source port, destination, destination port and protocol, attached to a subnet or a NIC. It works at L4 (addresses and ports, not URLs) and is stateful - replies to an allowed connection are allowed automatically.

Evaluation, per direction: lowest priority number first, first match wins. Your rules use 100-4096; below them sit default rules you cannot delete:

So isolating a subnet = specific allows at low numbers, then your own "deny VirtualNetwork" above the defaults (e.g. 4000). Use service tags (AzureLoadBalancer, Storage) instead of IP lists. Gotcha: az network nsg rule create without --destination-port-ranges defaults to 80. With NSGs on both subnet and NIC, both must allow.

Also asked: How would you isolate an application subnet so only the App Gateway subnet can reach it on 443? · What are service tags and why use them in NSG rules? · What can an NSG not filter, and what do you use instead?

Learn it: 23.3 Network security groups: priorities, defaults, service tags

How does Azure decide where a packet goes, and how do you send a subnet's internet traffic through a firewall? Mid

Every subnet has system routes: the VNet's range VnetLocal, peered ranges VNetPeering, 0.0.0.0/0 Internet, and the private ranges (10/8, 172.16/12, 192.168/16, 100.64/10) to None - dropped - unless something more specific covers them.

Choosing: longest prefix match first; on a tie, user-defined > BGP > system.

To force egress through a firewall: a route table with a UDR 0.0.0.0/0 -> VirtualAppliance <firewall private IP>, attached to the subnet (az network route-table route create ..., az network vnet subnet update --route-table). Traffic inside the VNet still matches the more specific VnetLocal and goes direct. A user 0/0 to an appliance also removes the private-range None routes.

Mistakes to avoid: asymmetric routing (out via the firewall, back direct - the stateful firewall drops it), routing AzureFirewallSubnet through itself, and breaking subnets that need direct internet (App Gateway v2). For a cluster behind it, the firewall must allow the endpoints the nodes need, or they never join. az network nic show-effective-route-table shows what really applies.

Also asked: What is a user-defined route and why would you use one? · Cluster nodes behind a firewall fail to join after creation. What would you check? · What is asymmetric routing and why does a firewall drop it?

Learn it: 23.7 Route tables: where packets go after the NSG says yes

After disabling public network access on a storage account or Key Vault, apps inside the VNet get 403. How do you diagnose it? Mid

Almost always DNS. A private endpoint gives the service a private IP in your subnet, but the app still uses the public name (stsysoplab.blob.core.windows.net), which must resolve to that IP from inside the VNet. Azure does it with a CNAME to *.privatelink.blob.core.windows.net, answered privately only where a matching private DNS zone is visible. Otherwise the name resolves to the public IP - which is now closed, hence a 403 that looks like auth.

  1. From inside the VNet: nslookup <name> - a public IP means the chain did not stop at the private zone.
  2. az network private-dns link vnet list ... - is the zone linked to this VNet?
  3. az network private-dns record-set a list ... - does the A record exist, and does it match the endpoint's IP? (A zone group keeps it in sync; hand-made records go stale.)
  4. Custom DNS servers must forward privatelink.* to Azure DNS 168.63.129.16.
  5. One endpoint per sub-resource: a blob endpoint does nothing for file.

Also asked: What is an Azure private endpoint? · What is the difference between a private endpoint and a service endpoint? · How would you organise private DNS zones across many subscriptions?

Learn it: 23.10 Private endpoints and the DNS that makes them work

Two spokes cannot communicate through the hub firewall. How do you troubleshoot it? Mid

Peering is not transitive, so spoke-to-spoke needs four pieces, and I check them in order:

  1. az network vnet peering list on both spokes and the hub - every peering Connected? Initiated = only one side exists, nothing flows.
  2. allowForwardedTraffic on the destination spoke's peering - traffic arriving from the firewall did not originate in the hub.
  3. Route table on the source subnet: destination spoke's prefix -> VirtualAppliance the firewall's IP. Without it the packet never goes to the hub.
  4. Route table on the destination subnet for the way back. Without it replies go nowhere, or around the firewall, which drops the asymmetric half-flow.
  5. The firewall rules allow the flow (its logs show allowed/denied).
  6. The NSGs on both subnets.

If the network is fine but a name is not, it is DNS: the private DNS zone must be linked to (or resolvable from) the calling spoke.

Also asked: Is VNet peering transitive, and what does that mean in practice? · What does gateway transit do in a hub-and-spoke network? · Why does a private endpoint in one spoke often not resolve from another spoke?

Learn it: 23.14 Hub and spoke: peering, transit, and why spokes cannot see each other

What is the difference between Azure Load Balancer and Application Gateway, and how does each typically fail? Mid

Both send traffic only to backends whose health probe passes.

Typical failures:

Also asked: Long-lived connections from pods hang after quiet periods. What is happening and how do you fix it? · Users get intermittent 502 errors from an Application Gateway. How do you investigate? · What is SNAT port exhaustion and how do you avoid it?

Learn it: 23.16 Load Balancer vs Application Gateway, and the idle timeout

What does Azure manage for you in AKS, and what do you manage? Mid

Azure runs the control plane: API server, etcd, scheduler, controller manager - in Microsoft's subscription. You see one endpoint (fqdn), never the machines; no etcd backups or certificate rotation of your own. The tier decides the SLA (Free: none; Standard: the production default).

You own the nodes: each node pool is a VM scale set you pay for, in the node resource group (MC_<rg>_<cluster>_<region>) together with the load balancer, public IPs, PVC disks and the kubelet identity. That group is AKS-managed - change it through az aks or Terraform, never by hand.

Also yours:

Also asked: How would you design node pools for a production AKS cluster? · How do you run AKS upgrades safely in production? · Why should you not change resources in the MC_ resource group by hand?

Learn it: 23.20 AKS from the Azure side

How would you size a subnet for an AKS cluster using Azure CNI? Mid

First ask which Azure CNI:

IPs = (nodes + surge) x (maxPods + 1)  + 5 reserved
50 nodes, maxPods 30, surge 1:  51 x 31 = 1,581 + 5 = 1,586
/22 = 1,019 usable  too small      /21 = 2,043 usable  fits

Count every pool in the subnet at its max (autoscaler maxCount), plus surge for upgrades - a subnet filled to the last node makes the next upgrade fail. Too small fails up front with InsufficientSubnetSize, and the message is the formula.

You cannot resize a used subnet or change a pool's subnet or maxPods, so the way out is a new subnet, a new pool, drain the old one - or Overlay on a new cluster. Also: the service CIDR (default 10.0.0.0/16) must not overlap your networks.

Also asked: Compare Azure CNI Overlay and Azure CNI with a node subnet. When would you choose each? · What happens when an AKS cluster runs out of IP addresses in its subnet? · Why must the cluster's service CIDR not overlap your corporate networks?

Learn it: 23.22 AKS networking: node subnet, overlay, and the IP bill

Pods are Pending but the cluster autoscaler isn't adding nodes. What do you check? Mid

The cluster autoscaler adds nodes only for pods that are Pending because nothing fits, judged by requests, not usage. It writes its reason as events on the pod: kubectl describe pod / kubectl get events.

And timing: a new node takes minutes to boot, so short spikes are the HPA's and headroom's job. Scale-down is blocked by PDBs, emptyDir pods, bare pods and safe-to-evict: "false".

Also asked: How does the cluster autoscaler decide to add or remove nodes? · How would you use spot node pools without risking availability? · What is the difference between the cluster autoscaler and the Horizontal Pod Autoscaler?

Learn it: 23.25 Cluster autoscaler, spot pools, and pool design

How do users get access to an Entra-integrated AKS cluster, and what does a developer need to deploy to one namespace? Mid

Two separate gates:

  1. ARM - may you get a kubeconfig? az aks get-credentials needs listClusterUserCredential, from Azure Kubernetes Service Cluster User Role. The kubeconfig holds no credential, only an exec plugin: kubelogin fetches an Entra token each time (kubelogin convert-kubeconfig -l azurecli reuses your az login). Missing it: AuthorizationFailed.
  2. Kubernetes API - may your token do this? Either Kubernetes RBAC with Entra group object ids as subjects, or Azure RBAC for Kubernetes (--enable-azure-rbac): role assignments like RBAC Reader/Writer/Admin, scoped to the cluster or to .../namespaces/shop. They are dataActions, so even a subscription Owner gets Forbidden ("User does not have access to the resource in Azure").

So a shop developer needs Cluster User Role + RBAC Writer on namespaces/shop, both granted to the team's group.

Then close the back door: --admin credentials bypass Entra and RBAC, so --disable-local-accounts - after setting up a break-glass group with Cluster Admin via PIM.

Also asked: Compare Kubernetes RBAC and Azure RBAC for Kubernetes authorization in AKS. · Why would you disable local accounts on an AKS cluster, and what must exist first? · How should automation authenticate to an AKS cluster?

Learn it: 23.31 AKS access: Entra ID, Azure RBAC for Kubernetes, local accounts

Practise these answers with flashcards and labs Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.