OnCallReady

Lesson 8.11 · Addressing & DNS · 22 min read

The routing table: which way does a packet go

In plain words

Picture a post office sorting letters. On the wall there is a list: "letters for this street go straight to the door", "letters for the north of the city go to the north depot", "everything else goes to the central office". When a letter could match two lines, the clerk uses the most specific one: "Oak Street, north district" beats "north district".

That list is the routing table. ip route shows it; default is "everything else"; via is the depot you hand the letter to; no via means you deliver it yourself on your own street. The most specific match wins (longest prefix), and ip route get IP asks the kernel which line it would actually use for one address.

Why this matters

"This box cannot reach the reports server, but everyone else can." The most common cause is that the box sends the traffic the wrong way. Linux decides the way from one table, and one command shows you its decision.

What you need to know already: Ch 1 · Network identity (1.9) - ip r showed default via 10.64.0.1, your Mac acting as the gateway. 8.3 - does an address fall inside a range.

A few words first

Every packet asks one question

"Which interface do I leave through, and do I hand the packet to a gateway or straight to the destination?" The routing table answers it. ip route (short: ip r) prints it:

$ ip route
default via 10.64.0.1 dev enp0s1 proto dhcp src 10.64.0.2 metric 100
10.64.0.0/24 dev enp0s1 proto kernel scope link src 10.64.0.2 metric 100
10.64.0.1 dev enp0s1 proto dhcp scope link src 10.64.0.2 metric 100

Each line starts with the destination range, then fields:

The second line exists because the interface has 10.64.0.2/24: the kernel knows everything in that /24 is on the same wire.

Longest prefix match

When several routes match an address, the most specific one wins - the one with the longest prefix. This rule is called longest prefix match. Metric only breaks ties between routes with the same prefix.

Imagine this table (wg0 is a VPN interface, 8.6):

10.0.0.0/8       via 10.8.0.1 dev wg0
10.20.0.0/16     via 10.8.0.1 dev wg0
10.20.5.0/24     via 10.64.0.254 dev enp0s1
default          via 10.64.0.1 dev enp0s1
10.20.5.9   -> matches /8, /16, /24 and default. /24 wins -> enp0s1 via .254
10.20.9.9   -> /8, /16, default. /16 wins           -> wg0
10.99.1.1   -> /8 and default. /8 wins               -> wg0
8.8.8.8     -> only default                          -> enp0s1 via .1

Every router on the internet decides this way, and so do cloud networks. "Why is traffic to that one subnet going somewhere odd?" is almost always a more specific route someone added.

Ask the kernel, do not guess: ip route get

ip route get ADDRESS prints the route the kernel would actually use for that one address:

$ ip route get 10.0.3.20
10.0.3.20 via 10.64.0.1 dev enp0s1 src 10.64.0.2 uid 1000
    cache

Read it as: to reach 10.0.3.20, hand it to 10.64.0.1, out of enp0s1, with source 10.64.0.2. (uid 1000 is your user id - the lookup was done for you; cache means the kernel remembered the result.) Use it whenever the box has more than one interface. Some variations:

$ ip route get 10.64.0.50
10.64.0.50 dev enp0s1 src 10.64.0.2 uid 1000
    cache

No via: on-link. The box will look for .50 directly on its own wire (how it does that is the next lesson).

$ ip route get 10.64.0.2
local 10.64.0.2 dev lo src 10.64.0.2 uid 1000
    cache <local>

Your own address goes over loopback (lo), never the wire.

# on a box with no default route (this one has one, so 10.9.9.9 goes via .1)
ip route get 10.9.9.9
RTNETLINK answers: Network is unreachable

RTNETLINK answers: is how the ip command reports an error from the kernel. "Network is unreachable" only happens when no route matches at all - no default route either. You meet it on isolated machines and on a box where someone deleted the default route.

What a missing route looks like to programs

ping sends a tiny "are you there?" message and prints the reply (it uses ICMP, the small protocol for network control messages; a protocol is an agreed format for talking). curl fetches a web page from the command line; -v (verbose) prints every step it takes.

# same box without a default route
ping -c1 10.9.9.9
ping: connect: Network is unreachable

curl -v http://10.9.9.9/
*   Trying 10.9.9.9:80...
* Immediate connect fail for 10.9.9.9: Network is unreachable
curl: (7) Failed to connect to 10.9.9.9 port 80 after 0 ms: Couldn't connect to server

(-c1 = send one message and stop.) Instant and local: no packet left the box. Compare that with a route that exists but goes the wrong way - say, a company range sent to the default gateway, which has no idea what to do with it. Then the packets leave, nothing comes back, and the program waits until it gives up (a timeout). The fastest way to tell the two apart is ip route get.

Adding and removing routes

Changing routes needs root:

# on a box with a second NIC, enp0s2 = 10.8.0.2/24 (the incident after this lesson has one)
ip route add 10.20.0.0/16 via 10.8.0.1 dev enp0s2
RTNETLINK answers: Operation not permitted
sudo ip route add 10.20.0.0/16 via 10.8.0.1 dev enp0s2
ip route get 10.20.1.10
10.20.1.10 via 10.8.0.1 dev enp0s2 src 10.8.0.2 uid 1000
    cache

(NIC = network interface card, the network port a machine plugs into; each shows up as an interface like enp0s1.) The errors you will hit:

RTNETLINK answers: File exists          that range already has a route; use
                                        'ip route replace' or delete it first
Error: Nexthop has invalid gateway.     the gateway is not on any network this
                                        box is directly on - it cannot reach it
Cannot find device "wg1"                typo, or the interface is not up yet
RTNETLINK answers: No such process      deleting a route that is not there

Nexthop has invalid gateway teaches something: a gateway must be directly reachable. You cannot route via 10.1.1.1 unless some interface is on a subnet that contains 10.1.1.1.

sudo ip route del 10.20.0.0/16
sudo ip route replace 10.20.0.0/16 via 10.8.0.1 dev enp0s2   # add or change

Routes added with ip are gone after a reboot

Same lesson as the swap fix in Ch 2 (swapoff plus /etc/fstab): the running state and the configuration are different things. On Ubuntu the network configuration is netplan - the YAML files in /etc/netplan/ you saw in Ch 1. (YAML is a config format where indentation with spaces shows nesting.)

# the same second-NIC box
cat /etc/netplan/60-corp.yaml
network:
  version: 2
  ethernets:
    enp0s2:
      addresses:
        - 10.8.0.2/24
      routes:
        - to: 10.20.0.0/16
          via: 10.8.0.1

Read it as: for interface enp0s2, give it address 10.8.0.2/24, and add a route to 10.20.0.0/16 via 10.8.0.1.

sudo netplan apply
ip route | grep 10.20
10.20.0.0/16 via 10.8.0.1 dev enp0s2 proto static

netplan apply turns the files into running state. proto static tells you the route came from configuration. Netplan is fussy: spaces only, never tabs, and it warns when a file is readable by everyone - chmod 600 (Ch 4), because these files can hold Wi-Fi and VPN passwords. sudo netplan try applies with an automatic rollback after 120 seconds unless you confirm. Use it on a remote box, where a wrong default route cuts off your own SSH session.

What you can now do

Why it helps

Routing questions come up constantly once there is more than one network: a VPN, a second NIC, a link to the office. The symptoms are distinctive once you know them. "Network is unreachable" instantly, before anything leaves the box, means no route. A slow timeout on one company range means a route sends it the wrong way. Traffic to one subnet mysteriously going through a firewall is a more specific route someone added.

With ip route get you answer "which way will this go" in one command instead of reading configs. The same longest-prefix rule runs cloud route tables and every router on the internet. You will also avoid the classic mistake of fixing a route live and losing it at reboot because netplan was never updated, and you will know to use netplan try before touching the default route over SSH.

Commands in this lesson

ip

FAQ

What decides which route wins: the prefix or the metric?

The prefix, first. The kernel picks the most specific matching route, the longest prefix. 10.20.5.0/24 beats 10.20.0.0/16, which beats 10.0.0.0/8, which beats default. Metric only breaks ties between routes with the same prefix, for example two default routes on two interfaces, where the lower metric wins.

Why use ip route get instead of reading ip route?

Because ip route get is the kernel's actual decision for one destination, including the chosen interface, next hop and source address. Reading ip route yourself means doing the longest-prefix match in your head, and you can miss a more specific route or extra rules the kernel applies. On any box with more than one interface or a VPN, ask the kernel.

What does "Nexthop has invalid gateway" mean?

The gateway you gave is not on any directly connected subnet, so the kernel has no way to reach it. A gateway must be one hop away: some interface must have an address in a subnet containing the gateway's IP. Fix the gateway address, or bring up the interface and address first. It is also why deleting an address can make routes via that subnet's gateway vanish.

Why did my route disappear after a reboot?

ip route add changes the running kernel state only, like swapoff without /etc/fstab in Ch 2. On Ubuntu the permanent configuration is netplan, in /etc/netplan/*.yaml under the interface's routes:. Apply with sudo netplan apply, or sudo netplan try on a remote box so a mistake rolls back automatically. Routes from configuration show proto static.

Is this the same as route on my Mac?

Same idea, different tools. macOS shows the table with netstat -rn and asks the kernel with route get IP, which is the equivalent of ip route get. The Linux route command and netstat come from net-tools, which Ubuntu Server no longer installs by default. Use the ip commands on Linux; they are what you will find on servers.

In an interview Junior

Explain this line: default via 10.64.0.1 dev enp0s1 proto dhcp src 10.64.0.2 metric 100

It is a route from ip route:

When several routes match, longest prefix match decides: the most specific route wins, and default loses to everything. I would not guess: ip route get IP prints the route the kernel actually picks. A route added with ip route add is gone after a reboot; the permanent one goes in netplan.

Also asked: What is the difference between "Network is unreachable" and a timeout? · How do you add a static route, and how do you make it survive a reboot? · What does longest prefix match mean?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.