Load balancers: ALB, NLB, health checks and their status codes
Users never talk to your instances; they talk to a load balancer, and when something breaks they see the load balancer's error page. Knowing what each status code means coming from the load balancer is half of debugging a web outage on AWS. This lesson builds that, then gives the load balancer a name with Route 53.
Need to know: an ALB (Application Load Balancer) works at HTTP: listeners on ports, rules by host/path/header, target groups of instances, IPs or Lambdas, health checks by HTTP status. An NLB (Network Load Balancer) works at TCP/UDP: static IPs per AZ, the client's source IP kept, millions of connections. From an ALB, 502 = the target refused or broke the connection, 503 = no target to send to, 504 = the target did not answer in time. With all targets unhealthy an ALB fails open and sends traffic to all of them anyway.
The objects
| Object | What it holds |
|---|---|
| load balancer | nodes in two or more subnets (one per AZ), a DNS name, security groups (ALB; NLBs optionally) |
| listener | protocol + port (HTTP 80, HTTPS 443 with a certificate), a default action, rules |
| rule | conditions (path-pattern, host-header, http-header, source-ip...) -> an action: forward, redirect, fixed-response, authenticate |
| target group | protocol + port, target type (instance, ip, lambda), the health check, attributes (deregistration delay, stickiness) |
| target | an instance ID or an IP, optionally with its own port |
$ LB=$(aws elbv2 describe-load-balancers --names try-alb --query 'LoadBalancers[0].LoadBalancerArn' --output text)
$ aws elbv2 describe-load-balancers --load-balancer-arns $LB --query 'LoadBalancers[0].[Type,Scheme,State.Code,DNSName]' --output text
application internet-facing active try-alb-1491553409.eu-central-1.elb.amazonaws.com
$ aws elbv2 describe-listeners --load-balancer-arn $LB --query 'Listeners[].[Protocol,Port,DefaultActions[0].Type]' --output text
HTTP 80 forward
$ TG=$(aws elbv2 describe-target-groups --names try-web --query 'TargetGroups[0].TargetGroupArn' --output text)
$ aws elbv2 describe-target-groups --target-group-arns $TG --query 'TargetGroups[0].[Protocol,Port,TargetType,HealthCheckPath,HealthCheckIntervalSeconds,HealthyThresholdCount,UnhealthyThresholdCount,Matcher.HttpCode]' --output text
HTTP 8080 instance /healthz 10 2 2 200
Never put an ALB's IP addresses anywhere: they change as AWS scales the nodes. The DNS name is the address (and a Route 53 alias, below, the human name).
Health checks
The load balancer's nodes probe each target on the health check port and path every interval seconds. healthy threshold consecutive passes make a target healthy, unhealthy threshold failures make it unhealthy. Only healthy targets get traffic - unless none is healthy.
$ aws elbv2 describe-target-health --target-group-arn $TG --query 'TargetHealthDescriptions[].[Target.Id,Target.Port,TargetHealth.State,TargetHealth.Reason]' --output text
i-02c9cb943f3ba6a34 8080 healthy None
i-09b9311391163d63f 8080 healthy None
$ curl -s http://$(aws elbv2 describe-load-balancers --load-balancer-arns $LB --query 'LoadBalancers[0].DNSName' --output text)/
<!doctype html>
<title>shop</title>
<h1>shop-web 2.4.1</h1>
<p>served by ip-10-99-10-30</p>
When a target is not healthy, the reason says why - learn these:
| Reason | Means | Usually |
|---|---|---|
Elb.RegistrationInProgress / Elb.InitialHealthChecking | just registered | wait one interval x threshold |
Target.Timeout | the probe got no answer | a security group/NACL drops it, or the app hangs |
Target.FailedHealthChecks | connection failed (refused) | nothing listens on that port |
Target.ResponseCodeMismatch | an HTTP answer, wrong code ([404]) | wrong health check path, or the app's /healthz failing |
Target.NotInUse | the group is not used by any listener rule | a forgotten wiring step |
Target.InvalidState | the instance is stopped |
Make a target fail on purpose - point the health check at a path the app does not have:
$ aws elbv2 modify-target-group --target-group-arn $TG --health-check-path /nope --query 'TargetGroups[0].HealthCheckPath' --output text
/nope
$ sleep 30
$ aws elbv2 describe-target-health --target-group-arn $TG --query 'TargetHealthDescriptions[].[Target.Id,TargetHealth.State,TargetHealth.Reason,TargetHealth.Description]' --output text
i-02c9cb943f3ba6a34 unhealthy Target.ResponseCodeMismatch Health checks failed with these codes: [404]
i-09b9311391163d63f unhealthy Target.ResponseCodeMismatch Health checks failed with these codes: [404]
$ curl -s -o /dev/null -w '%{http_code}\n' http://$(aws elbv2 describe-load-balancers --load-balancer-arns $LB --query 'LoadBalancers[0].DNSName' --output text)/
200
$ aws elbv2 modify-target-group --target-group-arn $TG --health-check-path /healthz --query 'TargetGroups[0].HealthCheckPath' --output text
/healthz
Both targets unhealthy - and the site still answers 200. That is fail-open: with zero healthy targets the ALB routes to all of them, because the health check is more likely wrong than every server at once. It is why a bad health check path does not cause an outage, and why "all targets unhealthy" plus errors usually means the targets really are broken.
The status codes from an ALB
| Code | Body says | Meaning |
|---|---|---|
| 502 Bad Gateway | awselb/2.0 | the target refused the connection, reset it, or sent a malformed answer: wrong port, app crashed, keep-alive timeout shorter than the ALB's 60 s idle timeout |
| 503 Service Temporarily Unavailable | awselb/2.0 | no registered targets in the group (or the rule matched nothing to forward to) |
| 504 Gateway Time-out | awselb/2.0 | the target accepted nothing or answered nothing in time: a security group/NACL dropping packets, an overloaded app |
| 4xx/5xx with the app's own body | the app's | the ALB passed the app's answer on - look at the app, not the ALB |
NLB in one paragraph
An NLB forwards TCP/UDP/TLS without looking inside: no paths, no HTTP codes, health checks by TCP connect (or HTTP on a path). Each AZ node gets a static IP (or your Elastic IP) - what partners put into their firewalls - and targets see the client's real IP. Use it for non-HTTP protocols, static addresses, extreme connection counts, or in front of an ALB when you need both. Security groups on NLBs are optional (since 2023); without one, the targets' groups must allow the clients.
Route 53: a name for it
Route 53 hosts DNS zones. For an AWS load balancer you create an alias record: Route 53 answers with the load balancer's current addresses, follows them as they change, costs nothing per query, and works at the zone apex (example.com itself), where a CNAME is not allowed. With EvaluateTargetHealth the record is considered unhealthy when the ALB has no healthy targets - the building block of failover routing (a primary and a secondary record, by health check).
$ aws route53 create-hosted-zone --name try.oncall-lab.example --caller-reference try-$(date +%s) --query 'DelegationSet.NameServers' --output text
ns-147.awsdns-32.org ns-1826.awsdns-07.co.uk ns-645.awsdns-53.com ns-55.awsdns-15.net
The four name servers are what the parent zone must delegate to. A record change is a change batch (JSON), applied atomically, PENDING until every Route 53 server has it (INSYNC) - the lab does it end to end.
In an interview: "What is the difference between 502, 503 and 504 from an ALB?" - 502: the target refused or broke the connection (wrong port, crash); 503: no target to send to; 504: the target did not answer in time (blocked by a security group/NACL, or too slow).
What you can do now
- read a load balancer's listeners, rules, target groups and target health;
- map a status code and a health reason to the layer that is broken;
- explain fail-open, and choose between ALB and NLB;
- point a name at a load balancer with an alias record.