OnCallReady

Lesson 23.16 · Azure II: Networking & AKS · 13 min read

Load Balancer vs Application Gateway, and the idle timeout

In plain words

Imagine two kinds of helpers at a busy entrance. The first one just waves each person towards one of several doors, without asking who they are or what they want; it's fast and simple. The second one opens each letter, reads the address, checks it isn't dangerous, and takes it to the right office. And both check each door regularly: if a door's porter doesn't answer, nobody gets sent there.

Azure Load Balancer is the first helper: layer 4, it spreads TCP and UDP connections by IP and port, and every type: LoadBalancer Service in AKS uses it. Application Gateway is the second: layer 7, it terminates TLS, routes by host and path, and adds a WAF. Both use health probes. The Load Balancer also forgets a connection after 4 minutes of silence, the idle timeout.

Users reach the shop through one public address, and behind it sit several pods. Something has to take each connection or request and pick a healthy backend. Azure has two products for it, working at different layers, and each has its own way of failing: a 502 from one, a silently dropped idle connection from the other.

What you need to know already: 9.1 (TCP, RST), 9.21 (HTTP), 9.23 (reverse proxies and load balancers), 9.15 (TLS termination), 16.6 (type: LoadBalancer Services), 16.19 (Ingress), 23.3 (NSGs).

Two products, two layers

Azure Load Balancer works at L4: it forwards TCP/UDP connections by address and port. Application Gateway (App Gateway, AGW) is a managed reverse proxy at L7 (the HTTP layer): it reads each request's host and path. It can also run a WAF (web application firewall: rules that block common attacks like SQL injection, from the OWASP list - a community catalogue of web attacks).

Azure Load BalancerApplication Gateway
layerL4 (TCP/UDP)L7 (HTTP/HTTPS)
seesIPs and portshosts, paths, headers, cookies
TLSpasses it throughterminates it (and can re-encrypt)
WAFnoyes (WAF_v2 SKU, OWASP rules)
routing5-tuple hashpath/host based, rewrites, redirects
in AKSevery type: LoadBalancer ServiceApplication Gateway Ingress Controller / App Gateway for Containers

5-tuple hash: the LB picks a backend from source IP, source port, destination IP, destination port and protocol, so one connection always lands on the same backend. AGIC (Application Gateway Ingress Controller) and App Gateway for Containers turn Kubernetes Ingress objects into App Gateway configuration.

Front Door is Azure's global L7 edge in front of either: a CDN (content delivery network - copies of your content served from locations near users), a WAF, and anycast (one IP announced from many locations, so users reach the nearest).

Health probes decide who gets traffic

Both products probe backends (send a test request every few seconds); an unhealthy backend gets nothing. For App Gateway, backend health is the first thing to read on any 502. az network application-gateway show-backend-health -g <rg> -n <gateway> prints each backend server and why it is (un)healthy:

$ az network application-gateway show-backend-health -g rg-oncall-lab -n agw-sysop --query "backendAddressPools[0].backendHttpSettingsCollection[0].servers[]"
[
  {
    "address": "10.20.0.200",
    "health": "Unhealthy",
    "healthProbeLog": "Received invalid status code: 404 in the backend server's HTTP response. As per the health probe configuration, 200-399 is the acceptable status code. Either modify probe configuration or resolve backend issues."
  }
]

All backends unhealthy = App Gateway answers 502 Bad Gateway to every client, even though the backends may be serving traffic perfectly well to anyone who asks the right path. Probe path, host header, port and expected status codes all have to match what the backend really answers. The incident in this chapter is one of these.

The 4-minute idle timeout

Chapter 9 met this from the TCP side. Here is the Azure side:

AKS (23.20) enables TCP reset on the load balancer it manages (named kubernetes, in the MC_ node resource group), and uses a 30-minute idle timeout on its outbound rule (the rule that carries pods' traffic out to the internet; SNAT, below). Inbound rules for your Services default to 4 minutes, changeable per Service with an annotation:

$ az network lb rule list -g MC_rg-oncall-lab_aks-sysop_westeurope --lb-name kubernetes --query "[].{name:name, port:frontendPort, idle:idleTimeoutInMinutes, reset:enableTcpReset}" -o table
Name                                      Port  Idle  Reset
----------------------------------------  ----  ----  -----
a3f1c0d2e4b54c6f8a9b0c1d2e3f4a5b-TCP-443  443   4     True
a3f1c0d2e4b54c6f8a9b0c1d2e3f4a5b-TCP-80   80    4     True

$ az network lb outbound-rule list -g MC_rg-oncall-lab_aks-sysop_westeurope --lb-name kubernetes -o table
AllocatedOutboundPorts  EnableTcpReset  IdleTimeoutInMinutes  Name             Protocol
----------------------  --------------  --------------------  ---------------  --------
0                       True            30                    aksOutboundRule  All
apiVersion: v1
kind: Service
metadata:
  name: orders-grpc
  annotations:
    service.beta.kubernetes.io/azure-load-balancer-tcp-idle-timeout: "30"
spec:
  type: LoadBalancer
  ...

Do not edit the AKS load balancer directly (az network lb rule update): the cloud provider reconciles it from the Service definitions and your change disappears at the next reconcile. Change the Service annotation, or az aks update --load-balancer-idle-timeout for the outbound rule.

The durable fix is on the client: TCP keepalives below the idle timeout (for example every 2-3 minutes), or application-level pings on long-lived connections (gRPC keepalive, database pool validation/idle eviction below 4 minutes). Then the timeout never triggers, whatever the infrastructure does.

SNAT, briefly

SNAT (source network address translation) rewrites the private source address of an outgoing connection to a public one, the way your home router does. Outbound connections from pods and VMs without public IPs leave through the load balancer's frontend (public) IPs, each with ~64,000 SNAT ports shared across the backend pool (the set of VMs behind it). Lots of short outbound connections to the same destination (connection-per-request HTTP clients, no pooling) exhaust them, and new connections fail intermittently. Symptoms look random; the metric is "SNAT connection count / failed" in the LB's metrics. Fixes: connection pooling (reusing connections instead of opening one per request), more outbound IPs, explicit allocated ports, or a NAT Gateway (a dedicated Azure resource for outbound SNAT with many more ports, attached to a subnet).

Reading an App Gateway configuration

$ az network application-gateway show -g rg-oncall-lab -n agw-sysop --query "{sku:sku.name, listeners:httpListeners[].{name:name, host:hostName, proto:protocol}, probes:probes[].{name:name, path:path}, rules:requestRoutingRules[].name}"
{
  "listeners": [ { "host": "shop.sysoplab.dev", "name": "https-shop", "proto": "Https" } ],
  "probes": [ { "name": "healthz", "path": "/healthz" } ],
  "rules": [ "shop" ],
  "sku": "WAF_v2"
}

The chain is listener (frontend IP, port, host, certificate) -> rule (which listener goes to which backend) -> backend pool (IPs or FQDNs) + HTTP settings (port, protocol, timeout, host override) + probe. A 502 is almost always in the right half of that chain.

Status codes and who sent them

502 from App Gateway    no healthy backend, or the backend reset/closed the connection
504 from App Gateway    backend accepted but did not answer within requestTimeout (default 30s)
403 from App Gateway    WAF blocked it (check the WAF logs, rule id and matched value)
502/503 from nginx      the ingress behind the gateway had no ready endpoints

Look at the response headers and body (curl -v, 9.21): App Gateway's own error pages say "Microsoft-Azure-Application-Gateway/v2" in the Server header. Know which box produced the error before you debug the one behind it.

What you can now do

Why it helps

Two of the most confusing production symptoms come from here. A 502 from Application Gateway when the backend is actually fine: the health probe uses the wrong path or host header, every backend is marked unhealthy, and the gateway answers 502 to everyone. And long-lived connections, database pools, gRPC streams, that hang after a quiet period because the load balancer's 4-minute idle timeout forgot them.

Knowing the layers lets you work out which box produced an error before debugging the one behind it, from the status code and the Server header. You'll also know not to edit the AKS-managed load balancer, which gets reconciled, but to use Service annotations, and to fix idle timeouts on the client with keepalives. SNAT exhaustion, the random outbound failures, is another classic you'll recognise.

Commands in this lesson

az

FAQ

When should I use Load Balancer versus Application Gateway?

Load Balancer for layer 4: any TCP or UDP service, very high throughput, no need to look inside the traffic, like every Kubernetes LoadBalancer Service or a database endpoint. Application Gateway for layer 7 HTTP and HTTPS: TLS termination, routing by host name or path, header rewrites, and a WAF against OWASP attacks. Many platforms use both: App Gateway or Front Door at the edge, and internal load balancers behind it.

Why does App Gateway return 502 when my app works?

Because App Gateway only sends traffic to backends its health probe considers healthy, and when none are, it answers 502 itself. The probe's path, host header, port and accepted status codes must match what the backend really serves: a probe to / on an app that returns 404 there, or a probe with the wrong host header for a virtual-hosted backend, marks everything unhealthy. show-backend-health gives the exact reason.

What happens when the idle timeout expires?

The load balancer forgets the flow. Without TCP reset, the next packet on that connection is silently dropped; the client thinks the connection is still open and its next request hangs until a TCP timeout. With TCP reset enabled, which AKS does, both ends get a RST and the client can reconnect quickly. The durable fix is client-side keepalives or pool eviction below the timeout, typically every 2-3 minutes.

How do I change the idle timeout for an AKS service?

With an annotation on the Service: service.beta.kubernetes.io/azure-load-balancer-tcp-idle-timeout: "30" sets 30 minutes for its inbound rules. For the outbound rule used by pods' egress, az aks update --load-balancer-idle-timeout. Don't edit the load balancer in the MC_ resource group directly: the cloud provider reconciles it from the Service definitions and your change is reverted at the next sync.

What is SNAT exhaustion?

Pods and VMs without their own public IPs share the load balancer's frontend IPs for outbound connections, each with about 64,000 source ports shared across the backend pool. Many short-lived connections to the same destination, typically HTTP clients without connection pooling, use them up, and new outbound connections fail intermittently with symptoms that look random. Fixes are connection pooling and reuse, more outbound IPs, explicitly allocated ports, or a NAT Gateway.

In an interview Mid

What is the difference between Azure Load Balancer and Application Gateway, and how does each typically fail?

Both send traffic only to backends whose health probe passes.

Typical failures:

Also asked: Long-lived connections from pods hang after quiet periods. What is happening and how do you fix it? · Users get intermittent 502 errors from an Application Gateway. How do you investigate? · What is SNAT port exhaustion and how do you avoid it?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.