Why this matters
"The app cannot reach X" is the ticket you will get most. Without a method, people jump to "it's the firewall" and an hour disappears while the name was resolving to the wrong address all along. This lesson puts the whole chapter into five checks, in order.
What you need to know already: this whole chapter - getent (8.16), ip route get (8.11), ip neigh (8.14), tcpdump (8.14). Ports from Ch 3.
Five questions, in order
Walk up the stack one question at a time and stop at the first one whose answer is wrong:
1. NAME what address does the program get for X? getent hosts X
2. ROUTE which way will the packet leave? ip route get IP
3. NEIGHBOUR does the next hop answer on the wire? ip neigh
4. PORT does anything answer on the port? nc -zv IP PORT
5. PACKETS what actually goes out, and what comes back? tcpdump host IP
The order matters because each step assumes the previous one is fine. Checking the firewall (steps 4-5) when the name resolves to the wrong address (step 1) is how an hour disappears.
1. Name
# a box with a stale /etc/hosts entry for api.lab (yours resolves it correctly)
getent hosts api.lab
10.0.99.99 api.lab
dig +short @10.0.3.53 api.lab
10.0.3.20
Different answers: stop here. The problem is on this box (/etc/hosts, nsswitch, resolved's settings or cache), not the network. Same answer: carry the IP forward. No answer from getent at all (exit code 2): check resolvectl status, the search list, and whether the name exists (dig).
2. Route
$ ip route get 10.0.3.20
10.0.3.20 via 10.64.0.1 dev enp0s1 src 10.64.0.2 uid 1000
cache
Three things to check against what you expect: the device, the next hop (or on-link), and the source address. The source matters on a box with several NICs: the far end's firewall allows specific sources, and replies go back to whatever source you used.
Failures here look like:
RTNETLINK answers: Network is unreachable no route at all
10.0.3.20 dev enp0s2 src 10.8.0.2 uid 1000 on-link through the WRONG NIC
10.20.1.10 via 10.64.0.1 dev enp0s1 ... the default route took a range
that should go elsewhere
The middle one is the wrong-mask bug from 8.14, and it is worth recognising on sight. Someone configured a second NIC as 10.8.0.2/8 instead of /24:
# the wrong-mask box (this chapter's last incident builds it)
ip -br a
lo UNKNOWN 127.0.0.1/8
enp0s1 UP 10.64.0.2/24
enp0s2 UP 10.8.0.2/8
ip route
default via 10.64.0.1 dev enp0s1 proto dhcp src 10.64.0.2 metric 100
10.0.0.0/8 dev enp0s2 proto kernel scope link src 10.8.0.2
...
The kernel now believes every 10.x address is on the enp0s2 wire (that /8 route beats the default route by longest prefix), so it stops sending them to any gateway and ARPs for them directly. The ones really on that segment work; everything else in 10.x breaks. It looks like a random part of the network went down.
3. Neighbour
For an on-link destination, or for the gateway:
# the wrong-mask box
ip neigh show dev enp0s2
10.8.0.1 dev enp0s2 lladdr 52:54:00:8c:01:01 REACHABLE
10.0.3.20 dev enp0s2 FAILED
FAILED for an address that has no business being on this segment is the wrong-mask bug confirmed from the other side. For the gateway itself, FAILED means the router is down or you are on the wrong segment, and nothing beyond it can work.
# the wrong-mask box
ping -c 2 10.0.3.20
PING 10.0.3.20 (10.0.3.20) 56(84) bytes of data.
From 10.8.0.2 icmp_seq=1 Destination Host Unreachable
From 10.8.0.2 icmp_seq=2 Destination Host Unreachable
Note the source in the message: From 10.8.0.2 - the enp0s2 address. The box is telling you which interface it tried.
4. Port
Most services talk TCP (Transmission Control Protocol): before any data flows, the two sides set up a connection with a quick three-message handshake. The first message is a SYN ("can we talk?"). If nothing listens on that port, the far machine answers with a RST ("no one here").
nc (netcat) opens a TCP connection and reports what happened. Flags: -z just test, send no data; -v say what happened; -w 3 give up after 3 seconds.
$ nc -zv -w 3 10.0.3.20 443
Connection to 10.0.3.20 443 port [tcp/https] succeeded!
$ nc -zv -w 3 10.0.3.12 5432
nc: connect to 10.0.3.12 port 5432 (tcp) timed out: Operation now in progress
$ nc -zv oncall-lab 9999
nc: connect to oncall-lab (127.0.1.1) port 9999 (tcp) failed: Connection refused
# and on the wrong-mask box:
nc -zv 10.0.3.20 443
nc: connect to 10.0.3.20 port 443 (tcp) failed: No route to host
([tcp/https] is nc naming the usual service on port 443.) Four answers, four places to look:
- succeeded - open.
- timed out - no reply at all: something dropped the SYN (a firewall), or the route loses it.
- Connection refused - the host answered with a RST: it is reachable, but nothing listens on that port.
- No route to host - step 3 surfacing: ARP failed, or a router sent back "host unreachable".
Later (Ch 9): the next chapter opens up TCP, the handshake and these failures in detail.
5. Packets
When the other steps do not settle it, look at the wire. Filters: host IP only packets to or from that address; and port 5432 only that port; -i any capture on every interface:
$ sudo tcpdump -nn -i any -c 3 host 10.0.3.12 and port 5432 & sleep 1; nc -zv -w 5 10.0.3.12 5432; wait
listening on any, link-type LINUX_SLL2 (Linux cooked v2), snapshot length 262144 bytes
10:40:01.101200 enp0s1 Out IP 10.64.0.2.41822 > 10.0.3.12.5432: Flags [S], seq 771820, ...
10:40:02.101237 enp0s1 Out IP 10.64.0.2.41822 > 10.0.3.12.5432: Flags [S], seq 771820, ...
10:40:04.101274 enp0s1 Out IP 10.64.0.2.41822 > 10.0.3.12.5432: Flags [S], seq 771820, ...
nc: connect to 10.0.3.12 port 5432 (tcp) timed out: Operation now in progress
Each line: time, interface, direction (Out), sender.port > receiver.port, and Flags [S] = a SYN. The same SYN three times, one, two, four seconds apart (the kernel resending it - a retransmit), and nothing back. The packets leave the right interface; they die somewhere after this box. That is a statement you can take to the network team, with timestamps.
-i any adds the interface and direction columns, which is exactly what you want when routing is the question: it shows which NIC the packet used.
Fixing it properly
For the wrong mask, the running fix and the permanent fix are different, as always. ip addr del removes an address from an interface, ip addr add puts one on:
# on the wrong-mask box - the incident at the end of this chapter
sudo ip addr del 10.8.0.2/8 dev enp0s2
sudo ip addr add 10.8.0.2/24 dev enp0s2
ip route | grep enp0s2
10.8.0.0/24 dev enp0s2 proto kernel scope link src 10.8.0.2
Deleting the address also deletes its automatic route, and any route whose gateway is no longer reachable - so a 10.20.0.0/16 via 10.8.0.1 route disappears with it and has to come back. Then correct the source of the mistake, the netplan file, and netplan apply, so a reboot does not bring it back.
Name, route, neighbour, port, packets - and write down the answer to each before moving on, because the incident summary is exactly that list.
What you can now do
- Debug "cannot reach X" in a fixed order, with one command per step.
- Tell open, timed out, refused and no route apart with
nc. - Spot and fix the wrong-mask bug, live and in netplan.