OnCallReady

Lesson 30.19 · AWS II: VPC, EC2, ELB & EKS · 14 min read

Load balancers: ALB, NLB, health checks and their status codes

In plain words

A load balancer is a receptionist in front of a team of clerks. Visitors only ever talk to the receptionist, who checks every few seconds which clerks are awake (health checks) and sends each visitor to one of them. If a clerk refuses to answer the receptionist says "502", if there is no clerk at all "503", and if a clerk is too slow "504".

Route 53 is the phone book entry that points people at the receptionist.

Load balancers: ALB, NLB, health checks and their status codes

Users never talk to your instances; they talk to a load balancer, and when something breaks they see the load balancer's error page. Knowing what each status code means coming from the load balancer is half of debugging a web outage on AWS. This lesson builds that, then gives the load balancer a name with Route 53.

Need to know: an ALB (Application Load Balancer) works at HTTP: listeners on ports, rules by host/path/header, target groups of instances, IPs or Lambdas, health checks by HTTP status. An NLB (Network Load Balancer) works at TCP/UDP: static IPs per AZ, the client's source IP kept, millions of connections. From an ALB, 502 = the target refused or broke the connection, 503 = no target to send to, 504 = the target did not answer in time. With all targets unhealthy an ALB fails open and sends traffic to all of them anyway.

The objects

ObjectWhat it holds
load balancernodes in two or more subnets (one per AZ), a DNS name, security groups (ALB; NLBs optionally)
listenerprotocol + port (HTTP 80, HTTPS 443 with a certificate), a default action, rules
ruleconditions (path-pattern, host-header, http-header, source-ip...) -> an action: forward, redirect, fixed-response, authenticate
target groupprotocol + port, target type (instance, ip, lambda), the health check, attributes (deregistration delay, stickiness)
targetan instance ID or an IP, optionally with its own port
$ LB=$(aws elbv2 describe-load-balancers --names try-alb --query 'LoadBalancers[0].LoadBalancerArn' --output text)
$ aws elbv2 describe-load-balancers --load-balancer-arns $LB --query 'LoadBalancers[0].[Type,Scheme,State.Code,DNSName]' --output text
application	internet-facing	active	try-alb-1491553409.eu-central-1.elb.amazonaws.com
$ aws elbv2 describe-listeners --load-balancer-arn $LB --query 'Listeners[].[Protocol,Port,DefaultActions[0].Type]' --output text
HTTP	80	forward
$ TG=$(aws elbv2 describe-target-groups --names try-web --query 'TargetGroups[0].TargetGroupArn' --output text)
$ aws elbv2 describe-target-groups --target-group-arns $TG --query 'TargetGroups[0].[Protocol,Port,TargetType,HealthCheckPath,HealthCheckIntervalSeconds,HealthyThresholdCount,UnhealthyThresholdCount,Matcher.HttpCode]' --output text
HTTP	8080	instance	/healthz	10	2	2	200

Never put an ALB's IP addresses anywhere: they change as AWS scales the nodes. The DNS name is the address (and a Route 53 alias, below, the human name).

Health checks

The load balancer's nodes probe each target on the health check port and path every interval seconds. healthy threshold consecutive passes make a target healthy, unhealthy threshold failures make it unhealthy. Only healthy targets get traffic - unless none is healthy.

$ aws elbv2 describe-target-health --target-group-arn $TG --query 'TargetHealthDescriptions[].[Target.Id,Target.Port,TargetHealth.State,TargetHealth.Reason]' --output text
i-02c9cb943f3ba6a34	8080	healthy	None
i-09b9311391163d63f	8080	healthy	None
$ curl -s http://$(aws elbv2 describe-load-balancers --load-balancer-arns $LB --query 'LoadBalancers[0].DNSName' --output text)/
<!doctype html>
<title>shop</title>
<h1>shop-web 2.4.1</h1>
<p>served by ip-10-99-10-30</p>

When a target is not healthy, the reason says why - learn these:

ReasonMeansUsually
Elb.RegistrationInProgress / Elb.InitialHealthCheckingjust registeredwait one interval x threshold
Target.Timeoutthe probe got no answera security group/NACL drops it, or the app hangs
Target.FailedHealthChecksconnection failed (refused)nothing listens on that port
Target.ResponseCodeMismatchan HTTP answer, wrong code ([404])wrong health check path, or the app's /healthz failing
Target.NotInUsethe group is not used by any listener rulea forgotten wiring step
Target.InvalidStatethe instance is stopped

Make a target fail on purpose - point the health check at a path the app does not have:

$ aws elbv2 modify-target-group --target-group-arn $TG --health-check-path /nope --query 'TargetGroups[0].HealthCheckPath' --output text
/nope
$ sleep 30
$ aws elbv2 describe-target-health --target-group-arn $TG --query 'TargetHealthDescriptions[].[Target.Id,TargetHealth.State,TargetHealth.Reason,TargetHealth.Description]' --output text
i-02c9cb943f3ba6a34	unhealthy	Target.ResponseCodeMismatch	Health checks failed with these codes: [404]
i-09b9311391163d63f	unhealthy	Target.ResponseCodeMismatch	Health checks failed with these codes: [404]
$ curl -s -o /dev/null -w '%{http_code}\n' http://$(aws elbv2 describe-load-balancers --load-balancer-arns $LB --query 'LoadBalancers[0].DNSName' --output text)/
200
$ aws elbv2 modify-target-group --target-group-arn $TG --health-check-path /healthz --query 'TargetGroups[0].HealthCheckPath' --output text
/healthz

Both targets unhealthy - and the site still answers 200. That is fail-open: with zero healthy targets the ALB routes to all of them, because the health check is more likely wrong than every server at once. It is why a bad health check path does not cause an outage, and why "all targets unhealthy" plus errors usually means the targets really are broken.

The status codes from an ALB

CodeBody saysMeaning
502 Bad Gatewayawselb/2.0the target refused the connection, reset it, or sent a malformed answer: wrong port, app crashed, keep-alive timeout shorter than the ALB's 60 s idle timeout
503 Service Temporarily Unavailableawselb/2.0no registered targets in the group (or the rule matched nothing to forward to)
504 Gateway Time-outawselb/2.0the target accepted nothing or answered nothing in time: a security group/NACL dropping packets, an overloaded app
4xx/5xx with the app's own bodythe app'sthe ALB passed the app's answer on - look at the app, not the ALB

NLB in one paragraph

An NLB forwards TCP/UDP/TLS without looking inside: no paths, no HTTP codes, health checks by TCP connect (or HTTP on a path). Each AZ node gets a static IP (or your Elastic IP) - what partners put into their firewalls - and targets see the client's real IP. Use it for non-HTTP protocols, static addresses, extreme connection counts, or in front of an ALB when you need both. Security groups on NLBs are optional (since 2023); without one, the targets' groups must allow the clients.

Route 53: a name for it

Route 53 hosts DNS zones. For an AWS load balancer you create an alias record: Route 53 answers with the load balancer's current addresses, follows them as they change, costs nothing per query, and works at the zone apex (example.com itself), where a CNAME is not allowed. With EvaluateTargetHealth the record is considered unhealthy when the ALB has no healthy targets - the building block of failover routing (a primary and a secondary record, by health check).

$ aws route53 create-hosted-zone --name try.oncall-lab.example --caller-reference try-$(date +%s) --query 'DelegationSet.NameServers' --output text
ns-147.awsdns-32.org	ns-1826.awsdns-07.co.uk	ns-645.awsdns-53.com	ns-55.awsdns-15.net

The four name servers are what the parent zone must delegate to. A record change is a change batch (JSON), applied atomically, PENDING until every Route 53 server has it (INSYNC) - the lab does it end to end.

In an interview: "What is the difference between 502, 503 and 504 from an ALB?" - 502: the target refused or broke the connection (wrong port, crash); 503: no target to send to; 504: the target did not answer in time (blocked by a security group/NACL, or too slow).

What you can do now

Why it helps

During an outage users see the load balancer's error, not your app's. Knowing what 502, 503 and 504 mean from an ALB, how health checks and their reason codes work, why a target group with every target unhealthy still serves traffic, and how alias records point a name at a load balancer turns a vague "the site is broken" into a specific broken layer within minutes.

Commands in this lesson

aws curl sleep

FAQ

When do I use an NLB instead of an ALB?

For anything that is not HTTP: TCP, UDP, TLS passthrough; when partners need static IP addresses for their allow-lists (one per AZ, Elastic IPs possible); when targets must see the client's real IP; or for very high connection rates. ALBs are for HTTP features: host and path rules, redirects, fixed responses, authentication, WAF.

Why are my targets unhealthy with Target.Timeout?

The health check got no answer at all: something drops the packets. Usually the targets' security group does not allow the health check port from the load balancer's security group, or a network ACL blocks the request or its reply. FailedHealthChecks without a code means the connection was refused, so nothing listens on that port.

Why does the site work although every target is unhealthy?

With no healthy targets an ALB fails open and routes to all of them, assuming the health check is wrong rather than every server. That is why a wrong health check path does not cause an outage on its own. If errors appear at the same time, the targets really are broken.

What is the deregistration delay?

When a target is deregistered (scale-in, deploy), the load balancer stops sending it new requests but lets open ones finish for this long, 300 seconds by default. Shorten it for fast deploys of short requests; keep it long for long downloads or websockets. It is AWS's version of connection draining.

Why use an alias record instead of a CNAME?

Route 53 resolves an alias to the load balancer's current addresses itself, so it works at the zone apex where a CNAME is forbidden, costs nothing per query, and can follow the target's health. Load balancer IPs change, so never put them into records or firewalls.

In an interview Mid

What is the difference between 502, 503 and 504 from an ALB?

502 Bad Gateway: the ALB reached a target that refused or broke the connection or sent an invalid response - a wrong port, a crashed app, a keep-alive timeout shorter than the ALB's. 503: there was no target to send to - no registered targets, or no rule forwarding. 504: the target did not answer in time - a security group or network ACL dropping packets, or an overloaded app. describe-target-health and its reason codes show which layer.

Also asked: How do ALB health checks decide a target is healthy? · When would you choose an NLB over an ALB? · What does fail-open mean for a target group?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.