Why this matters
One of the strangest tickets you will get: "the connection works, small requests work, but big downloads hang forever". Nothing is down. Packets above a certain size are silently vanishing somewhere on the path. Once you know the pattern you can prove it with two pings.
What you need to know already: routes and interfaces (8.11), ping and ICMP (Chapter 8, 9.11), the mss option in the SYN (9.1), netplan (8.11).
The sizes
Every packet carries a header (addresses, flags - the envelope) in front of its data. Sizes are in bytes.
MTU the largest IP packet an interface will send (Ethernet: 1500)
MSS the largest TCP payload in one segment (MTU - 20 IP - 20 TCP = 1460)
MTU = Maximum Transmission Unit. MSS = Maximum Segment Size. ip link show prints an interface's MTU:
$ ip link show enp0s1
2: enp0s1: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc fq_codel state UP mode DEFAULT group default qlen 1000
link/ether 3e:a1:5c:77:2b:09 brd ff:ff:ff:ff:ff:ff
(UP = enabled, mtu 1500, link/ether = its MAC address from 8.14; the rest you can ignore for now.)
Each side announces its MSS in the SYN (options [mss 1460,...] in tcpdump) and both use the smaller. That only protects the two ends' own links. Anything in between with a smaller MTU is a different problem.
When the path is smaller than the ends
A tunnel carries packets inside other packets: a VPN (virtual private network) wraps your packet in an encrypted outer packet to cross the internet to another site. The outer header takes room, so the payload a tunnel can carry shrinks:
VXLAN (virtual networks inside data centres) 50 bytes -> inner 1450
WireGuard (a modern VPN) 60 (v4) / 80 (v6)
IPsec (site-to-site VPN) 50-73, depends on cipher
GRE 24
PPPoE (home DSL) 8 -> 1492
A 1500-byte packet reaches a VPN gateway whose tunnel only carries 1400. A router could fragment it - cut it into smaller packets - but TCP asks it not to. Two things can happen:
- The packet has DF (Don't Fragment) set - and modern TCP always sets it. The router drops it and sends back ICMP "fragmentation needed" (type 3, code 4) with the MTU it can carry. The sender lowers its path MTU (the smallest MTU on the whole way) for that destination and resends smaller. This is Path MTU Discovery (PMTUD), and it works invisibly.
- Somewhere between, a firewall blocks all ICMP "for security". The router still drops the packet, but its message never arrives. The sender keeps sending 1500-byte packets that keep vanishing. An MTU black hole.
What a black hole looks like
The signature is unmistakable once you know it: small things work, big things hang.
# the MTU black-hole lab (next mission): reports.lab behind a tunnel on enp0s2
nc -zv reports.lab 443 # handshake: small packets
Connection to reports.lab (10.20.1.10) 443 port [tcp/https] succeeded!
curl -s http://reports.lab/health # a tiny response
OK
curl -s -m 20 http://reports.lab/export # a large response
curl: (28) Operation timed out after 20001 milliseconds with 0 out of 184320 bytes received
curl -v -m 20 https://reports.lab/ # TLS: the certificate flight is big
* Trying 10.20.1.10:443...
* Connected to reports.lab (10.20.1.10) port 443
* ALPN: curl offers h2,http/1.1
* TLSv1.3 (OUT), TLS handshake, Client hello (1):
... nothing, until the timeout ...
The TLS case catches everyone. TLS (9.15) opens with its own handshake: the client's first message (ClientHello) is small, but the server's reply carries its certificates - several full-size segments - and never arrives. "TLS hangs after Client hello" is an MTU problem far more often than a TLS one.
Measure it: ping with DF
-M do sets Don't Fragment; -s sets the ICMP payload size; -c 1 sends one ping. The IP packet is payload + 8 (ICMP header) + 20 (IP header):
-s 1472 -> 1472 + 28 = 1500 bytes
-s 1372 -> 1400
# in the black-hole lab
ping -M do -s 1473 -c 1 10.20.1.10
PING 10.20.1.10 (10.20.1.10) 1473(1501) bytes of data.
ping: local error: message too long, mtu=1500
That one never left: bigger than our own interface.
# in the black-hole lab
ping -M do -s 1472 -c 2 10.20.1.10 # exactly 1500
PING 10.20.1.10 (10.20.1.10) 1472(1500) bytes of data.
--- 10.20.1.10 ping statistics ---
2 packets transmitted, 0 received, 100% packet loss, time 1001ms
ping -M do -s 1372 -c 2 10.20.1.10 # 1400
PING 10.20.1.10 (10.20.1.10) 1372(1400) bytes of data.
1380 bytes from 10.20.1.10: icmp_seq=1 ttl=61 time=3.21 ms
1380 bytes from 10.20.1.10: icmp_seq=2 ttl=61 time=3.25 ms
Big pings vanish silently, small ones get through: black hole, and you have bracketed the path MTU. Halve the gap until you find the exact size. If ICMP were allowed, the big ping would have told you directly:
From 10.8.0.1 icmp_seq=1 Frag needed and DF set (mtu = 1400)
tracepath
# in the black-hole lab
tracepath -n 10.20.1.10
1?: [LOCALHOST] pmtu 1500
1: 10.8.0.1 0.500ms
1: 10.8.0.1 0.400ms
2: 10.20.1.10 3.200ms reached
Resume: pmtu 1500 hops 2 back 2
tracepath is a traceroute cousin (-n: no name lookups): it walks the path router by router (each router is a hop), discovers the path MTU and prints pmtu 1400 where it drops - if the ICMP comes back. In a black hole it reports 1500 all the way, which is itself the clue: the ping test says 1400 fits and 1500 does not, and nobody told tracepath.
The fixes
Lower the MTU on the interface that leads into the small path:
# in the black-hole lab (enp0s2 is the tunnel NIC there)
sudo ip link set dev enp0s2 mtu 1400
ip link show enp0s2 | head -1
3: enp0s2: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1400 qdisc fq_codel state UP mode DEFAULT group default qlen 1000
TCP now announces MSS 1360 and never sends a segment that is too big. Persist it in netplan (mtu: 1400 under the interface). It affects everything on that interface, which is fine for a dedicated corp/VPN NIC and not fine for your only NIC.
Per route: ip route add 10.20.0.0/16 via 10.8.0.1 dev enp0s2 mtu 1400 limits only that destination.
ip link set dev X mtu N changes it now (lost at reboot); netplan's mtu: key keeps it.
MSS clamping on the router or firewall in the middle rewrites the MSS option in passing SYNs so both ends agree on segments that fit - the standard fix on VPN gateways, and what the Linux firewall tool iptables does with TCPMSS --clamp-mss-to-pmtu.
Stop blocking ICMP type 3. "Fragmentation needed" is not an attack; filtering it breaks TCP. Allow ICMP destination-unreachable through every firewall.
net.ipv4.tcp_mtu_probing=1 makes Linux probe for a working segment size after it detects a black hole - a useful safety net on clients, not a fix.
Where you will meet this
- Site-to-site VPNs between a cloud network and the company's own data centre: a new route through the VPN, and "only large file downloads fail".
- Any overlay network (VXLAN and friends) whose MTU was set to 1500 by mistake: traffic between machines hangs on big responses only.
- A firewall someone "hardened" by blocking all ICMP.
Later (Ch 15): Kubernetes puts every app on an overlay network like this, so a wrong MTU there looks exactly like this lesson.
What you can now do
- Recognise the pattern "small works, big hangs" as a path MTU problem.
- Measure the path MTU with
ping -M do -s SIZE. - Fix it on a host (
ip link set ... mtu, netplan) and name the real fixes upstream: MSS clamping and letting ICMP through.