OnCallReady

Lesson 9.23 · TCP, TLS & HTTP · 22 min read

Reverse proxies, load balancers and forward proxies

In plain words

A reverse proxy is like a receptionist at an office. Visitors never walk to the engineers' desks; they talk to the receptionist, who walks back, asks the right person, and comes back with the answer. If the engineer isn't at their desk, the receptionist says so (502). If the whole team is out, "nobody available" (503). If the engineer takes too long to answer, the receptionist gives up (504). If the visitor walks out before the answer arrives, the receptionist writes that down too (499).

nginx in front of the orders app is exactly that receptionist: two separate conversations, client to nginx and nginx to the app on 127.0.0.1:8080, and every 5xx is nginx telling you what happened on the second one.

Why this matters

Almost no web app is reached directly. There is a proxy in front - nginx on this box - and when users see 502, 503 or 504, the code is the proxy telling you what went wrong behind it. Read the code and the proxy's log line and you know which component to look at, before anyone opens the app's code.

What you need to know already: HTTP requests, headers and status codes (9.21), refused vs timeout (9.1), nginx as a systemd service and its logs in /var/log/nginx/ (Chapter 7 read them with grep and awk), TLS (9.15).

A reverse proxy is two connections

A proxy is a program that receives your request and makes it on your behalf. A reverse proxy sits in front of servers (clients think it is the server); a forward proxy sits in front of clients (the end of this lesson). The server the proxy forwards to is its upstream.

client --(1)--> nginx --(2)--> upstream (the app)

The client never talks to the app. nginx accepts connection (1), opens connection (2), and translates. Every 5xx a proxy returns describes what happened on (2), and nginx writes the reason to its error.log.

An nginx site is a server { } block: listen the port, server_name the hostname it answers for, and location /api/ { } what to do with paths starting /api/ - here proxy_pass (forward to the app), a timeout, and headers to add:

server {
    listen 80;
    server_name orders.lab;
    location /api/ {
        proxy_pass http://127.0.0.1:8080;
        proxy_read_timeout 5s;
        proxy_set_header Host              $host;
        proxy_set_header X-Forwarded-For   $proxy_add_x_forwarded_for;
        proxy_set_header X-Forwarded-Proto $scheme;
    }
}

502, 503, 504, 499 - and the log line for each

502 Bad Gateway - the proxy could not get a valid response from the upstream: nothing listening, the connection was reset, or it answered garbage.

# the orders.lab proxy lab in this chapter, with orders stopped
curl -s -o /dev/null -w '%{http_code}\n' http://orders.lab/api/orders
502
sudo tail -1 /var/log/nginx/error.log
2026/09/23 10:12:01 [error] 906#906: *1843 connect() failed (111: Connection refused) while connecting to upstream, client: 127.0.0.1, server: orders.lab, request: "GET /api/orders HTTP/1.1", upstream: "http://127.0.0.1:8080/api/orders", host: "orders.lab"

The line: date, level [error], nginx's process ids, a request number (*1843), the reason, then the client, the site, the request and the upstream it tried. 111: Connection refused - errno 111 (9.8), the upstream is down. Other 502 lines: upstream prematurely closed connection while reading response header (the app crashed or closed mid-request), no live upstreams while connecting to upstream (every server in the upstream block is marked failed).

504 Gateway Timeout - the upstream was reached and did not answer in time:

2026/09/23 10:14:11 [error] 906#906: *1851 upstream timed out (110: Connection timed out) while reading response header from upstream, client: 127.0.0.1, server: orders.lab, request: "GET /api/payments/refund HTTP/1.1", upstream: "http://127.0.0.1:8080/api/payments/refund", host: "orders.lab"

proxy_read_timeout (default 60s) is how long nginx waits for the response header. "while connecting" instead of "while reading" means the connect itself timed out (proxy_connect_timeout) - a network problem, not a slow app.

503 Service Unavailable - no backend available to try at all. Load balancers return it when every backend fails its health check (a regular test request, below). Also: nginx's own rate limiting (limit_req, which caps requests per second) answers 503 by default, not 429, unless you set limit_req_status 429.

499 (nginx only) - the client gave up and closed the connection before nginx had a response. You see it in access.log (the one-line-per-request log you parsed with awk in 7.8), not error.log:

127.0.0.1 - - [23/Sep/2026:10:15:02 +0000] "GET /api/payments/refund HTTP/1.1" 499 0 "-" "curl/8.14.1"

A wall of 499s means your clients' timeouts are shorter than your latency. The client is not broken; you are slow.

The triage in one line: 503 - look at health checks and endpoints; 504 - look at the upstream's latency and what it waits on; 502 - look at the upstream process and the connection to it; 499 - look at your own latency.

X-Forwarded-For and X-Forwarded-Proto

Behind a proxy the app sees the proxy's address, not the client's, and plain HTTP even when the client used HTTPS: TLS ended at the proxy (TLS termination - the proxy holds the certificate and decrypts). The proxy says what it saw in headers:

X-Forwarded-For:   203.0.113.24, 10.0.4.17     client, then each proxy it passed
X-Forwarded-Proto: https                       the scheme the CLIENT used
X-Real-IP:         203.0.113.24                nginx convention, single value

Two rules: only trust these from your own proxies (anyone can send X-Forwarded-For: 127.0.0.1 - configure the app or set_real_ip_from with the proxy's addresses), and read the left-most untrusted entry, not blindly the first.

The redirect loop

An app that enforces HTTPS ("if the request is not https, redirect to https") behind a proxy that terminates TLS and does not send X-Forwarded-Proto:

client --https--> nginx --http--> app    "that was http, go to https://shop.lab/"
client --https--> nginx --http--> app    "that was http, go to https://shop.lab/"
...
# the shop.lab lab in this chapter, before the fix:
curl -sL -o /dev/null https://shop.lab/
curl: (47) Maximum (50) redirects followed
curl -sI https://shop.lab/ | grep -i location
location: https://shop.lab/

A redirect to the same URL you requested is the signature. Fix it in the proxy (proxy_set_header X-Forwarded-Proto $scheme; - $scheme is nginx's variable for "http" or "https") and tell the app to trust it (every web framework has a "trust proxy headers" setting).

L4 vs L7

The L-numbers are layers of the networking model you met in Chapter 8: layer 4 is TCP (connections and ports), layer 7 is the application protocol (HTTP).

L4 (TCP)   forwards connections. Fast, protocol-blind. Cannot see paths, headers
           or status codes; cannot retry a failed HTTP request; one long-lived
           HTTP/2 connection = one backend forever.
           Most cloud "network load balancers" are L4.
L7 (HTTP)  terminates the connection, parses HTTP. Routes by host and path,
           terminates TLS, adds headers, retries idempotent requests, health
           checks on a real URL, balances per request.
           nginx is L7; so are cloud "application gateways".

"Why is this 500 not retried?" - because the thing in front is L4 and never saw a

  1. "Why does one backend get all the traffic from one client?" - because an L4

balancer picked it once, when the connection opened.

Health checks: active - the LB requests /health on a schedule and stops sending traffic to backends that fail; passive - the LB notices real requests failing and ejects the backend (nginx max_fails/fail_timeout). Use both: active catches dead backends before users do, passive catches the ones that pass /health and fail real work.

Algorithms (how the LB picks a backend): round robin (each in turn - the default everywhere), least connections (better when request cost varies), consistent hashing (the same key - user, cache key - always goes to the same backend, and adding or removing a backend moves few keys).

Later (Ch 16): Kubernetes has both kinds built in - an L4 balancer for every service and L7 "ingress" proxies that are often nginx itself - so everything in this lesson carries over.

Forward proxies: HTTP_PROXY and friends

In a bank, outbound traffic to the internet goes through a forward proxy (Squid is a common one). Tools find it through environment variables (2.21), and they disagree on the details:

http_proxy     for http:// URLs. curl reads ONLY the lowercase one (the
               uppercase HTTP_PROXY is ignored on purpose - "httpoxy").
https_proxy    for https:// URLs. curl accepts HTTPS_PROXY too.
no_proxy       comma-separated exceptions: hosts, domain suffixes (.lab),
               IPs, CIDRs (curl 7.86+). NO_PROXY also accepted.
$ https_proxy=http://proxy.lab:3128 curl -v -o /dev/null https://api.lab/
* Uses proxy env variable https_proxy == 'http://proxy.lab:3128'
*   Trying 10.0.3.30:3128...
* Connected to proxy.lab (10.0.3.30) port 3128
* CONNECT tunnel: HTTP/1.1 negotiated
> CONNECT api.lab:443 HTTP/1.1
> Host: api.lab:443
> User-Agent: curl/8.14.1
> Proxy-Connection: Keep-Alive
>
< HTTP/1.1 503 Service Unavailable
< Server: squid/6.13
< X-Squid-Error: ERR_DNS_FAIL 0
<
* CONNECT tunnel failed, response 503
curl: (56) CONNECT tunnel failed, response 503

For HTTPS the client asks the proxy to open a tunnel (CONNECT host:443) and does TLS through it. The proxy sits outside the internal DNS view (split horizon, 8.22), cannot resolve api.lab, and refuses. Internal names must be in no_proxy. The usual traps:

What you can now do

Why it helps

Every web request in a modern platform goes through at least one proxy, usually several: a cloud gateway at the edge, then nginx or a similar reverse proxy, then the app. When users see 502s, 503s or 504s, the code alone tells you where to look, and nginx's error.log tells you why: "Connection refused" (the app is down), "no live upstreams", "upstream timed out" (the app waits on something).

You'll also meet the classics: an infinite HTTPS redirect loop because X-Forwarded-Proto isn't set, rate limiting that answers 503 instead of 429, one client's traffic stuck on one backend behind an L4 balancer, and on a corporate network, no_proxy missing the internal domain so every internal call goes to Squid and fails.

FAQ

What's the difference between 502 and 504?

502 Bad Gateway means the proxy could not get a valid response: the upstream refused the connection, reset it, closed it early or sent garbage. Look at the upstream process. 504 Gateway Timeout means the upstream accepted the request and did not send response headers within proxy_read_timeout (60 seconds by default in nginx). Look at what the app is waiting on. If the error log says "while connecting", the connect itself timed out, which is a network problem.

What does 499 mean? It's not in the HTTP spec.

It is nginx-specific: the client closed the connection before nginx had a response to send. It appears in access.log, not error.log, and the response size is 0. A few are normal (users closing tabs). A wall of 499s means clients have shorter timeouts than your latency: typically an upstream service or load balancer gave up waiting. The client is not broken; you are slow.

Why does my app see the proxy's IP instead of the client's?

Because the app's TCP connection really comes from the proxy. The proxy reports what it saw in headers: X-Forwarded-For (client, then each proxy), X-Forwarded-Proto (the scheme the client used) and nginx's X-Real-IP. Configure the app or nginx's set_real_ip_from to trust these headers only from your own proxies, since any client can send a fake X-Forwarded-For: 127.0.0.1.

Is a cloud "network load balancer" the same as nginx in front of an app?

No. A network load balancer is L4: it forwards TCP or UDP connections and never sees paths, headers or status codes, so it cannot route by URL or retry a failed request. nginx as a reverse proxy is L7: it terminates HTTP, routes by host and path, terminates TLS and balances per request. They are often stacked: an L4 balancer spreading connections over several nginx servers, which then route each request.

I exported https_proxy in /etc/environment. Why doesn't my service use it?

Because /etc/environment is read by pam_env at login, so it affects new login sessions only, not your current shell and never systemd services. A service gets its environment from its unit: Environment= or EnvironmentFile= in a drop-in. Also, curl only reads lowercase http_proxy, and Java ignores these variables entirely, needing -Dhttps.proxyHost, -Dhttps.proxyPort and -Dhttp.nonProxyHosts.

In an interview Junior

Users get 502s from nginx in front of your application. How do you troubleshoot?

A 502 means nginx could not get a valid response on its connection to the upstream, so look behind it:

  1. sudo tail /var/log/nginx/error.log - the reason is in the line: connect() failed (111: Connection refused) = the app is down or on another port; upstream prematurely closed connection = it crashed or closed mid-request; no live upstreams = every backend marked failed.
  2. Check the app: systemctl status, sudo ss -tlnp for the port it listens on, its logs.
  3. curl the upstream directly (curl -v http://127.0.0.1:8080/...) to skip nginx.

Compare: 504 = upstream reached but slow (proxy_read_timeout, 60 s default); 503 = no backend available; 499 in access.log = the client gave up first.

Also asked: What is the difference between a forward proxy and a reverse proxy? · What is X-Forwarded-For, and why should you not trust it blindly? · What is the difference between an L4 and an L7 load balancer?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.