Why this matters
"Service A cannot connect to service B" is one of the most common tickets you will get. Without a method, people restart things and hard-code IPs. With one, you can name the broken layer in a few minutes - usually from the error text alone.
What you need to know already: DNS lookups and
dig(8.16, 8.18); the method for "I cannot reach X" on a host (8.29); refused vs timed out vs reset (9.1);ss -tlnp(9.5);curl -v(9.21); container networks, published ports and the 127.0.0.1 bind (11.15);nsenterand borrowing tools for distroless images (10.38).
"It can't connect" is not a diagnosis. Every connection goes through four steps, and each one fails with its own signature - the same order as 8.29, now for containers. Test them in order, from inside the caller's network namespace, and the answer is usually obvious by step 3.
1. name does the name resolve, to the address you expect?
2. route is that address reachable from here (same network)?
3. port is anything listening on that address and port?
4. app does the thing listening answer correctly?
Where to stand
Test from where the caller stands, not from the host. The host has different DNS, different routes, and sees published ports that containers do not use.
Most application images have no tools at all (distroless, slim - 10.38). Do not install tools into them - borrow the namespace:
# checkout, payments, fx, stock: the containers of the mission below
docker run --rm -it --network container:checkout nicolaka/netshoot
checkout ~ dig +short payments
checkout ~ nc -zv payments 8080
nicolaka/netshoot is an image packed with network tools (dig, nc, curl, ss, tcpdump). --network container:checkout puts the netshoot container in checkout's network namespace: same interfaces, same IP, same /etc/resolv.conf, same localhost. Whatever it sees, checkout sees. Nothing is changed in the target. (-it gives you an interactive shell in it; the lines after the prompt are commands typed there.)
The same trick from the host, with the host's own tools:
PID=$(docker inspect -f '{{.State.Pid}}' checkout)
sudo nsenter -t "$PID" -n ss -tlnp
sudo nsenter -t "$PID" -n curl -sv http://payments:8080/
nsenter runs a command inside another process's namespaces: -t PID = whose, -n = only the network namespace, so you keep the host's filesystem and tools. Careful: curl from the host inside the container's net namespace still reads the host's /etc/resolv.conf, so names resolve differently. For DNS questions, use netshoot.
Step 1: the name
docker run --rm --network container:checkout nicolaka/netshoot dig payments
;; ->>HEADER<<- opcode: QUERY, status: NXDOMAIN, id: 1147
;; SERVER: 127.0.0.11#53(127.0.0.11) (UDP)
dig NAME asks DNS and prints the whole answer (8.18). status: NXDOMAIN means "no such name"; SERVER is who answered. NXDOMAIN from 127.0.0.11: Docker's embedded DNS (11.15) does not know that name on the networks checkout is attached to. Either the target is on another network (user-defined networks are isolated from each other), the caller is on the default bridge (no DNS at all), or the name is wrong. Compare:
docker inspect -f '{{range $k, $v := .NetworkSettings.Networks}}{{$k}} {{end}}' checkout payments
shop
pay-net
(range $k, $v := loops over the Networks map and prints each key - the network name. One line per container: checkout is on shop, payments on pay-net.)
Fix by attaching one side to the other's network (docker network connect), never by hard-coding an IP. Error signatures of a DNS failure from apps:
curl: (6) Could not resolve host: payments
wget: bad address 'payments'
nc: getaddrinfo for host "payments" port 8080: Name does not resolve
java.net.UnknownHostException: payments
dial tcp: lookup payments on 127.0.0.11:53: no such host (Go)
getaddrinfo ENOTFOUND payments (Node)
Steps 2 and 3: route and port
docker run --rm --network container:checkout nicolaka/netshoot nc -zv fx 9000
nc: connect to fx (172.20.0.3) port 9000 (tcp) failed: Connection refused
nc -zv HOST PORT (netcat) tries a TCP connection and reports: -z = just test, send no data; -v = say what happened. Here the name resolved (you see the IP) and the host answered - with a TCP RST, the "go away" packet from 9.1. Something is at that address, but nothing accepts on that port on that interface. Look at the listener from inside the target:
docker run --rm --network container:fx nicolaka/netshoot ss -tlnp
State Recv-Q Send-Q Local Address:Port Peer Address:PortProcess
LISTEN 0 511 127.0.0.1:9000 0.0.0.0:* users:(("app",pid=1,fd=6))
Read the Local Address:Port column (9.5). 127.0.0.1:9000: it listens on the container's own loopback, which nobody else can reach. The app needs to bind 0.0.0.0. The other common answer here is a different port (0.0.0.0:8081 when the caller uses 8080).
Refused vs timed out is the most useful distinction in networking:
Connection refused reached the host, nothing listening there -> port/bind problem
timed out packets went nowhere -> routing, firewall, a dead host, wrong IP
reset by peer accepted, then dropped -> a proxy with no upstream, TLS mismatch
(Between containers on the same Docker bridge you will mostly see "refused"; timeouts show up with firewalls and with isolated networks, as in 11.15.)
Step 4: the application
docker run --rm --network container:checkout nicolaka/netshoot curl -sv http://stock:8081/health
* Trying 172.20.0.4:8081...
* Connected to stock (172.20.0.4) port 8081
> GET /health HTTP/1.1
> Host: stock:8081
> User-Agent: curl/8.11.1
> Accept: */*
>
< HTTP/1.1 200 OK
< Content-Type: text/plain
curl -v (verbose, 9.21) shows every step in one go: Trying (the name resolved), Connected (the port is open), the request lines (>), and the response lines (<) starting with the status line. A 502/503 here comes from a proxy (like nginx, 9.23) that could not reach its upstream - the server behind it - so you start again at step 1, from the proxy's namespace.
The published-port trap
docker ps --format '{{.Names}}\t{{.Ports}}'
stock 0.0.0.0:18081->8081/tcp
(--format with \t prints a tab between fields.) 18081 exists on the host (docker-proxy and a NAT rule that forwards it, 11.15). Between containers the published port is irrelevant: http://stock:8081, always the container port. And localhost inside a container is that container.
A one-screen checklist
docker inspect -f '{{json .NetworkSettings.Networks}}' A B same network?
netshoot --network container:A dig +short B name?
netshoot --network container:A nc -zv B PORT port open?
netshoot --network container:B ss -tlnp listening where?
netshoot --network container:A curl -sv http://B:PORT/path app answers?
("netshoot --network container:A" is short for docker run --rm --network container:A nicolaka/netshoot.)
What you can now do
- debug "A cannot reach B" in four ordered steps, from A's point of view
- name the broken layer from the error text: DNS, route, port or app
- borrow tools into a container's network namespace without changing it