OnCallReady

Lesson 8.29 · Addressing & DNS · 19 min read

A method for "I cannot reach X"

In plain words

When the TV does not work, a good repair person does not open the back first. They check in order: is it plugged in, is the socket working, is the cable connected, is the right input selected, and only then look inside. Each check only makes sense if the one before passed.

"I cannot reach X" works the same way, with five checks: the name (getent hosts X), the route (ip route get IP), the neighbour (ip neigh), the port (nc -zv IP PORT) and the packets (tcpdump). You stop at the first one whose answer is wrong. On oncall-lab the chapter's incidents all fall to this order, including the wrong subnet mask that makes a random part of 10.x disappear.

Why this matters

"The app cannot reach X" is the ticket you will get most. Without a method, people jump to "it's the firewall" and an hour disappears while the name was resolving to the wrong address all along. This lesson puts the whole chapter into five checks, in order.

What you need to know already: this whole chapter - getent (8.16), ip route get (8.11), ip neigh (8.14), tcpdump (8.14). Ports from Ch 3.

Five questions, in order

Walk up the stack one question at a time and stop at the first one whose answer is wrong:

1. NAME       what address does the program get for X?        getent hosts X
2. ROUTE      which way will the packet leave?                  ip route get IP
3. NEIGHBOUR  does the next hop answer on the wire?             ip neigh
4. PORT       does anything answer on the port?                 nc -zv IP PORT
5. PACKETS    what actually goes out, and what comes back?      tcpdump host IP

The order matters because each step assumes the previous one is fine. Checking the firewall (steps 4-5) when the name resolves to the wrong address (step 1) is how an hour disappears.

1. Name

# a box with a stale /etc/hosts entry for api.lab (yours resolves it correctly)
getent hosts api.lab
10.0.99.99      api.lab
dig +short @10.0.3.53 api.lab
10.0.3.20

Different answers: stop here. The problem is on this box (/etc/hosts, nsswitch, resolved's settings or cache), not the network. Same answer: carry the IP forward. No answer from getent at all (exit code 2): check resolvectl status, the search list, and whether the name exists (dig).

2. Route

$ ip route get 10.0.3.20
10.0.3.20 via 10.64.0.1 dev enp0s1 src 10.64.0.2 uid 1000
    cache

Three things to check against what you expect: the device, the next hop (or on-link), and the source address. The source matters on a box with several NICs: the far end's firewall allows specific sources, and replies go back to whatever source you used.

Failures here look like:

RTNETLINK answers: Network is unreachable      no route at all
10.0.3.20 dev enp0s2 src 10.8.0.2 uid 1000     on-link through the WRONG NIC
10.20.1.10 via 10.64.0.1 dev enp0s1 ...        the default route took a range
                                               that should go elsewhere

The middle one is the wrong-mask bug from 8.14, and it is worth recognising on sight. Someone configured a second NIC as 10.8.0.2/8 instead of /24:

# the wrong-mask box (this chapter's last incident builds it)
ip -br a
lo               UNKNOWN        127.0.0.1/8
enp0s1           UP             10.64.0.2/24
enp0s2           UP             10.8.0.2/8
ip route
default via 10.64.0.1 dev enp0s1 proto dhcp src 10.64.0.2 metric 100
10.0.0.0/8 dev enp0s2 proto kernel scope link src 10.8.0.2
...

The kernel now believes every 10.x address is on the enp0s2 wire (that /8 route beats the default route by longest prefix), so it stops sending them to any gateway and ARPs for them directly. The ones really on that segment work; everything else in 10.x breaks. It looks like a random part of the network went down.

3. Neighbour

For an on-link destination, or for the gateway:

# the wrong-mask box
ip neigh show dev enp0s2
10.8.0.1 dev enp0s2 lladdr 52:54:00:8c:01:01 REACHABLE
10.0.3.20 dev enp0s2 FAILED

FAILED for an address that has no business being on this segment is the wrong-mask bug confirmed from the other side. For the gateway itself, FAILED means the router is down or you are on the wrong segment, and nothing beyond it can work.

# the wrong-mask box
ping -c 2 10.0.3.20
PING 10.0.3.20 (10.0.3.20) 56(84) bytes of data.
From 10.8.0.2 icmp_seq=1 Destination Host Unreachable
From 10.8.0.2 icmp_seq=2 Destination Host Unreachable

Note the source in the message: From 10.8.0.2 - the enp0s2 address. The box is telling you which interface it tried.

4. Port

Most services talk TCP (Transmission Control Protocol): before any data flows, the two sides set up a connection with a quick three-message handshake. The first message is a SYN ("can we talk?"). If nothing listens on that port, the far machine answers with a RST ("no one here").

nc (netcat) opens a TCP connection and reports what happened. Flags: -z just test, send no data; -v say what happened; -w 3 give up after 3 seconds.

$ nc -zv -w 3 10.0.3.20 443
Connection to 10.0.3.20 443 port [tcp/https] succeeded!
$ nc -zv -w 3 10.0.3.12 5432
nc: connect to 10.0.3.12 port 5432 (tcp) timed out: Operation now in progress
$ nc -zv oncall-lab 9999
nc: connect to oncall-lab (127.0.1.1) port 9999 (tcp) failed: Connection refused
# and on the wrong-mask box:
nc -zv 10.0.3.20 443
nc: connect to 10.0.3.20 port 443 (tcp) failed: No route to host

([tcp/https] is nc naming the usual service on port 443.) Four answers, four places to look:

Later (Ch 9): the next chapter opens up TCP, the handshake and these failures in detail.

5. Packets

When the other steps do not settle it, look at the wire. Filters: host IP only packets to or from that address; and port 5432 only that port; -i any capture on every interface:

$ sudo tcpdump -nn -i any -c 3 host 10.0.3.12 and port 5432 & sleep 1; nc -zv -w 5 10.0.3.12 5432; wait
listening on any, link-type LINUX_SLL2 (Linux cooked v2), snapshot length 262144 bytes
10:40:01.101200 enp0s1 Out IP 10.64.0.2.41822 > 10.0.3.12.5432: Flags [S], seq 771820, ...
10:40:02.101237 enp0s1 Out IP 10.64.0.2.41822 > 10.0.3.12.5432: Flags [S], seq 771820, ...
10:40:04.101274 enp0s1 Out IP 10.64.0.2.41822 > 10.0.3.12.5432: Flags [S], seq 771820, ...
nc: connect to 10.0.3.12 port 5432 (tcp) timed out: Operation now in progress

Each line: time, interface, direction (Out), sender.port > receiver.port, and Flags [S] = a SYN. The same SYN three times, one, two, four seconds apart (the kernel resending it - a retransmit), and nothing back. The packets leave the right interface; they die somewhere after this box. That is a statement you can take to the network team, with timestamps.

-i any adds the interface and direction columns, which is exactly what you want when routing is the question: it shows which NIC the packet used.

Fixing it properly

For the wrong mask, the running fix and the permanent fix are different, as always. ip addr del removes an address from an interface, ip addr add puts one on:

# on the wrong-mask box - the incident at the end of this chapter
sudo ip addr del 10.8.0.2/8 dev enp0s2
sudo ip addr add 10.8.0.2/24 dev enp0s2
ip route | grep enp0s2
10.8.0.0/24 dev enp0s2 proto kernel scope link src 10.8.0.2

Deleting the address also deletes its automatic route, and any route whose gateway is no longer reachable - so a 10.20.0.0/16 via 10.8.0.1 route disappears with it and has to come back. Then correct the source of the mistake, the netplan file, and netplan apply, so a reboot does not bring it back.

Name, route, neighbour, port, packets - and write down the answer to each before moving on, because the incident summary is exactly that list.

What you can now do

Why it helps

"The app cannot reach X" is the ticket you will get most as a platform engineer, from developers, from alerts and in incident channels. Without a method, people jump to the firewall or blame "the network", and an hour disappears while the name was resolving to a stale /etc/hosts entry all along.

The method gives you three things: speed (most issues die at step 1 or 2), evidence ("connection attempts leave enp0s1 at 10:40:01, 10:40:02, 10:40:04, nothing comes back" is something a network team can act on), and a ready-made incident summary, because your notes for each step are the timeline. And "how would you troubleshoot connectivity" is asked in nearly every SRE interview; a structured answer stands out.

Commands in this lesson

ip nc tcpdump

FAQ

Why start with the name and not with ping?

Because every later step uses the address the name resolves to. If the program gets 10.0.99.99 from a stale /etc/hosts entry, pinging the right IP proves nothing about the program's problem, and checking firewalls for the right IP wastes time. getent hosts X shows the address the program really uses. Carry that address forward into the route, neighbour and port checks.

What do timeout, refused and no route to host each tell me?

A timeout means packets were dropped somewhere with no reply: a firewall, or routing that loses the traffic. Connection refused means the host answered with a RST: it is reachable but nothing listens on that port. No route to host means ARP failed on the local segment or a router sent back "host unreachable". Three different places to look.

Why use tcpdump with -i any?

Because routing is often the question, and -i any adds the interface and direction to each line (enp0s1 Out). You see which NIC the packets actually leave by, which confirms or refutes what ip route get told you. Filter tightly (host IP and port 5432) so the output is readable, and note the timestamps: repeated SYNs at 1, 2 and 4 seconds are retransmits with no answer.

If I fix the address with ip addr, is it fixed?

Only until the next reboot or netplan apply. ip addr and ip route change running state; the configuration lives in /etc/netplan/. Also, deleting an address removes its automatic route and any route whose gateway is no longer reachable, so re-add those. Then correct the netplan file and apply it, so the mistake does not come back.

What if all five steps look fine but the app still fails?

Then the problem is above the network. If nc -zv succeeds, the connection works, so look at what happens next: the service's own logs (journalctl -u, Ch 2), its error messages, the request the program actually sends (curl -v), permissions or passwords, and limits like the number of open connections. The next chapter covers what happens after the connection is open.

In an interview Junior

What is the difference between "connection refused" and "connection timed out"?

Both come from trying to open a TCP connection, for example with nc -zv -w 3 IP PORT:

They fit into the order for any "cannot reach X": name (getent hosts), route (ip route get), neighbour (ip neigh), port (nc -zv), packets (tcpdump). With tcpdump a timeout shows as the same Flags [S] retransmitted with nothing back.

Also asked: An application cannot connect to a database. How do you troubleshoot it? · How do you check whether a port is open on a remote host? · How would you use tcpdump to see whether packets leave the box?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.