OnCallReady

NetworkingLinuxDockerSRE · 5 min read

nginx 502 Bad Gateway but the API is up: reading error.log

A 502 is nginx saying its own connection to the upstream failed. How to read error.log, tell refused from timed out, and test from where nginx stands.

Users get a 502. The API team says their service is healthy, and from their shell it is:

terminal
$ curl -si http://ordersgw.lab/api/orders
HTTP/1.1 502 Bad Gateway
Server: nginx/1.26.3 (Ubuntu)
Content-Type: text/html
Content-Length: 166

Both are true. The API can be up and nginx can still fail to reach it.

What a 502 actually means

nginx here is a reverse proxy: it takes the client's request, opens its own connection to the app behind it (the upstream), and passes the answer back. The status code describes what happened on that second connection, not on yours.

  • 502 Bad Gateway: nginx got nothing usable from the upstream. It could not connect, or the upstream closed the connection or sent garbage.
  • 504 Gateway Timeout: nginx reached the upstream and waited longer than its timeout.
  • 499 (nginx-only, access.log): the client hung up before nginx answered.

So the right question is: can nginx, from where it runs, connect to the address in proxy_pass? The answer is written down for you.

The diagnosis path

1. Confirm nginx is the one saying 502

The Server: nginx header says the proxy produced the error. With several proxies in a row (a cloud load balancer, then nginx), make sure you read the right one's logs.

2. Read error.log, not access.log

The access log only records the status. The error log records why:

terminal
$ tail -1 /var/log/nginx/error.log
2026/09/22 20:00:10 [error] 906#906: *1844 connect() failed (111: Connection refused) while connecting to upstream, upstream: "http://127.0.0.1:8080/api/orders", client: 10.64.0.2, server: ordersgw.lab, request: "GET /api/orders HTTP/1.1", host: "ordersgw.lab"

Three parts matter: the errno (111: Connection refused), the phase (while connecting to upstream), and the upstream it tried. Is that the address you expected? A lot of 502s end right here.

The lines you will meet most often:

error.log saysstatusmeaning
connect() failed (111: Connection refused) while connecting to upstream502nothing listens on that IP:port
upstream prematurely closed connection while reading response header from upstream502the app accepted, then died or closed mid-request
no live upstreams while connecting to upstream502every server in the upstream block is marked failed
upstream timed out (110: Connection timed out) while connecting to upstream504SYNs unanswered: firewall drop, wrong host, dead box
upstream timed out (110: Connection timed out) while reading response header from upstream504connected, but the app did not answer within proxy_read_timeout

Refused is fast and definite: the upstream's kernel answered with a reset because nothing is bound to that port. Timed out means silence. Different causes, different places to look.

3. Test from where nginx stands

Repeat nginx's connection yourself, to the exact address from the log, from the same machine:

terminal
$ curl -sv http://127.0.0.1:8080/api/orders
*   Trying 127.0.0.1:8080...
* connect to 127.0.0.1 port 8080 from 127.0.0.1 port 38625 failed: Connection refused
* Failed to connect to 127.0.0.1 port 8080 after 0 ms: Couldn't connect to server

Same result as nginx: the problem is the upstream, not the proxy.

4. Is anything listening, and is it the right thing?

terminal
$ sudo ss -ltnp 'sport = :8080'
State Recv-Q Send-Q Local Address:Port Peer Address:PortProcess

An empty table: nothing listens on 8080. Then ask the service manager why:

terminal
$ systemctl status orders --no-pager
○ orders.service - Orders API (Spring Boot)
     Active: inactive (dead) since Tue 2026-09-22 20:00:09 UTC; 400ms ago

Stopped, crashed and restarting in a loop, still starting (a JVM can take several seconds before it listens - 502 until then), or bound to a different address such as 127.0.0.1 when nginx calls the box's IP. ss shows the bound address in the Local Address column, so check it too.

The container version of the same 502

In Docker the API answers on localhost:18080 from the host, and nginx in its own container still returns 502. Test from inside nginx's network namespace with a throwaway debug container:

terminal
$ docker run --rm --network container:portal-web-1 nicolaka/netshoot curl -sS -m 3 http://api:18080/
curl: (7) Failed to connect to api port 18080 after 0 ms: Couldn't connect to server
$ docker run --rm --network container:portal-web-1 nicolaka/netshoot curl -sS -m 3 http://api:8080/
{"service":"portal","version":"2.0.3"}
$ docker port portal-api-1
8080/tcp -> 0.0.0.0:18080

18080 is the published port: a door on the host. Between containers you use the container port, 8080. The config said proxy_pass http://api:18080;, so nginx connected to a port nothing listens on, and got refused.

Two more container traps behind the same symptom:

  • proxy_pass http://localhost:8080 inside the nginx container means the nginx container itself, not the host and not the API.
  • Containers find each other by name only on a shared network. If the proxy and the API sit on different user-defined networks, api does not resolve from nginx at all. A proxy that fronts a backend network has to join it. (When every name stops resolving, the resolver itself is down: a different hunt.)

In the official nginx image, error.log goes to stderr: read it with docker logs.

The fix

Fix whichever link the log points at: start the upstream, correct the proxy_pass host and port, or put the proxy on the backend network. Then reload nginx (sudo nginx -t && sudo systemctl reload nginx, or recreate the container) and check the status code is 200 again:

terminal
$ curl -s -o /dev/null -w '%{http_code}\n' http://ordersgw.lab/api/orders
200

Keeping it from coming back

  • Health-check the upstream through the proxy path, not only on the app's own port. A check on :18080 from the host stays green while every user gets a 502.
  • Give the proxy several upstream servers in an upstream block, so one restart is not an outage. In Kubernetes the same idea is readiness probes and a preStop pause, so the proxy stops sending to a pod before it goes away: see zero-downtime rolling updates. An Ingress that returns 502 or 503 for every request often fronts a Service with no endpoints.
  • Keep timeouts ordered: each layer's timeout shorter than the layer in front of it, or the outer layer gives up first and the inner ones keep working for nobody.

Practise it

Two labs in the course: 502, 504 and 499, made to order (9.24) makes nginx produce each code on purpose so you learn the matching error.log lines, and Incident: nginx says 502, the API is up (11.19) is the Docker version with two faults to find.

OnCallReady is free, with no ads and no tracking. RSS · All posts