Why this matters
Two classic outages come from simple arithmetic on ports and queues. A client suddenly cannot reach one server ("Cannot assign requested address") while everything else works. A server that is visibly listening makes every client time out. Both look like "the network", and neither is.
What you need to know already: ports (Chapter 8), socket states and TIME-WAIT (9.5), the LISTEN queue columns in ss (9.5), sysctl (9.1).
A connection is a 4-tuple
A tuple is just a fixed group of values. The kernel tells connections apart by four of them:
(source IP, source port, destination IP, destination port)
10.64.0.2:45120 -> 10.0.3.20:443
Every connection must have a unique tuple. The destination is fixed by what you connect to; your source IP is usually fixed; so the only thing that varies is the source port, picked from the ephemeral (short-lived) range - ports the kernel lends to outgoing connections:
$ sysctl net.ipv4.ip_local_port_range
net.ipv4.ip_local_port_range = 32768 60999
28,232 ports. Per destination IP and port - connections to 10.0.3.20:443 and to 10.0.3.12:5432 do not compete.
How a client runs out
A client that opens a new connection for every request and closes it first leaves each one in TIME-WAIT for 60 seconds, holding its source port:
28,232 ports / 60 s = ~470 new connections per second, sustained, to ONE
destination - then connect() fails
# the port-exhaustion lab later in this chapter
ss -Htan state time-wait '( dport = :443 )' | wc -l
28229
curl -v https://api.lab/
* Trying 10.0.3.20:443...
* Immediate connect fail for 10.0.3.20: Cannot assign requested address
curl: (7) Failed to connect to api.lab port 443 after 0 ms: Couldn't connect to server
Cannot assign requested address is EADDRNOTAVAIL - the kernel's name for this error (an errno: the numbered error code a system call returns): no free source port for that destination. Java says it as java.net.NoRouteToHostException: Cannot assign requested address, which sends people looking at routing. Other destinations still work fine, which is the second clue.
The fix is connection reuse, not sysctls
HTTP keep-alive means "keep this connection open after the response, I will send another request on it". A connection pool is a small set of open connections an app keeps and lends to each request in turn. Together they send thousands of requests over a handful of connections. No new tuples, no TIME-WAIT pile, and no TCP + TLS handshake per request (which is also most of the latency for small requests). The sysctls people reach for:
net.ipv4.ip_local_port_range wider range: more headroom, same bug
net.ipv4.tcp_tw_reuse lets NEW outgoing connections reuse a TIME-WAIT
tuple when TCP timestamps make it safe. Default 2
= loopback only. Setting 1 is reasonable on
clients; it is a mitigation.
net.ipv4.tcp_tw_recycle removed from Linux in 4.12. It broke clients
behind NAT. If a blog post tells you to set it,
close the blog post.
Behind a NAT gateway (Chapter 8), the same arithmetic happens on the NAT's ports. SNAT (source NAT) rewrites your private source address and port into the gateway's public ones, so every machine behind it shares one pool of ports: SNAT port exhaustion. Cloud NAT gateways often hand each machine only a small fixed share. A fleet of apps that does not reuse connections exhausts it and outbound calls start failing now and then. Same fix: reuse connections. Then more public addresses on the gateway.
The server side: backlog and the accept queue
A listening socket has a queue of completed connections waiting for the application to accept() them (9.5). Its size is the backlog the application asked for, capped by the kernel setting net.core.somaxconn:
$ sysctl net.core.somaxconn
net.core.somaxconn = 4096
$ ss -ltn
State Recv-Q Send-Q Local Address:Port Peer Address:Port Process
LISTEN 0 511 0.0.0.0:80 0.0.0.0:*
LISTEN 0 100 0.0.0.0:8090 0.0.0.0:*
LISTEN 0 4096 *:22 *:*
For a LISTEN row, Send-Q is the backlog. nginx asks for 511; the Java web server inside the app on 8090 for 100 (its accept-count setting); and port 22 shows systemd's default for a socket unit, 4096 (sshd is socket-activated on Ubuntu - systemd listens and starts sshd on demand; started by hand it asks for 128).
A Java web app handles requests with a fixed set of worker threads (a thread is a line of work inside one process; each worker handles one request at a time). When every worker thread is busy (say, all blocked waiting on a database), nobody calls accept(). The queue fills:
$ ss -ltn '( sport = :8090 )'
State Recv-Q Send-Q Local Address:Port Peer Address:Port Process
LISTEN 101 100 0.0.0.0:8090 0.0.0.0:*
Recv-Q 101 on a queue of 100 - full. From then on the kernel drops new SYNs (or the final ACK) instead of refusing them, so clients see a timeout, not "refused", against a port that is visibly listening. The counters prove it:
$ nstat -az TcpExtListenOverflows TcpExtListenDrops
#kernel
TcpExtListenOverflows 1843 0.0
TcpExtListenDrops 1843 0.0
nstat prints the kernel's network counters: -a absolute totals (not the change since last run), -z include counters that are zero. Without -a it prints the change since you last ran it: run it twice a few seconds apart and see if it is still climbing. Raising the backlog only buys queue; the fix is whatever is holding the workers - usually a missing timeout on a downstream call (a call to a service further down the chain).
Keepalive, and the load balancer's idle timeout
TCP keepalive (not the same thing as HTTP keep-alive) sends an empty probe on an idle connection so both ends (and everything in between) know it is alive. Linux defaults are useless for this:
$ sysctl net.ipv4.tcp_keepalive_time net.ipv4.tcp_keepalive_intvl net.ipv4.tcp_keepalive_probes
net.ipv4.tcp_keepalive_time = 7200
net.ipv4.tcp_keepalive_intvl = 75
net.ipv4.tcp_keepalive_probes = 9
First probe after two hours of idle (7200 s), then every 75 s, 9 times - and only on sockets whose application enabled keepalive at all (the SO_KEEPALIVE socket option).
Load balancers, NAT gateways and firewalls remember every connection that passes through them in a table, and forget idle ones after an idle timeout - 4 minutes is a common cloud default. They do it silently. Neither end is told. The next time either side writes, it gets an RST - "Connection reset by peer" out of nowhere, on the first request after a quiet period. Database pools, message consumers, long streams: anything long-lived and quiet.
Fixes: keepalives below the timeout, set in the client or pool (for example a 2-minute keepalive, or a pool that checks or throws away idle connections); or raise the LB's idle timeout; or have the LB send an RST when it forgets a flow, so clients at least learn about it immediately. sudo sysctl -w net.ipv4.tcp_keepalive_time=120 (-w = write) changes the default for applications that enable keepalive; put it in a file under /etc/sysctl.d/ to keep it after a reboot.
ss -o shows it working:
ESTAB 0 0 10.64.0.2:45120 10.0.3.20:443 timer:(keepalive,1min48sec,0)
Later (Ch 22): Azure's load balancer is one of these devices, with a 4-minute default idle timeout and its own SNAT port limits - the arithmetic above is exactly what you will size there.
What you can now do
- Explain "Cannot assign requested address" as source ports used up towards one destination, and prove it with
ssandip_local_port_range. - Recognise a full accept queue (LISTEN Recv-Q above Send-Q) and look for what the workers are stuck on.
- Explain why an idle connection gets reset, and where the keepalive belongs.