OnCallReady

Lesson 9.8 · TCP, TLS & HTTP · 18 min read

Ports, queues and exhaustion

In plain words

Picture a car park outside a stadium with 28,000 numbered spaces. Every car that visits the stadium needs its own space. When a car leaves, its space stays roped off for one minute before anyone else may use it, just to be safe. If more than about 470 new cars arrive every second and each stays only briefly, the car park fills with roped-off spaces and the next car is turned away, even though the stadium is half empty.

The spaces are ephemeral source ports (ip_local_port_range), the rope is TIME-WAIT, and "turned away" is Cannot assign requested address. Carpooling, sending many requests over one kept-open connection, is the fix.

Why this matters

Two classic outages come from simple arithmetic on ports and queues. A client suddenly cannot reach one server ("Cannot assign requested address") while everything else works. A server that is visibly listening makes every client time out. Both look like "the network", and neither is.

What you need to know already: ports (Chapter 8), socket states and TIME-WAIT (9.5), the LISTEN queue columns in ss (9.5), sysctl (9.1).

A connection is a 4-tuple

A tuple is just a fixed group of values. The kernel tells connections apart by four of them:

(source IP, source port, destination IP, destination port)
10.64.0.2:45120  ->  10.0.3.20:443

Every connection must have a unique tuple. The destination is fixed by what you connect to; your source IP is usually fixed; so the only thing that varies is the source port, picked from the ephemeral (short-lived) range - ports the kernel lends to outgoing connections:

$ sysctl net.ipv4.ip_local_port_range
net.ipv4.ip_local_port_range = 32768	60999

28,232 ports. Per destination IP and port - connections to 10.0.3.20:443 and to 10.0.3.12:5432 do not compete.

How a client runs out

A client that opens a new connection for every request and closes it first leaves each one in TIME-WAIT for 60 seconds, holding its source port:

28,232 ports / 60 s  =  ~470 new connections per second, sustained, to ONE
                        destination - then connect() fails
# the port-exhaustion lab later in this chapter
ss -Htan state time-wait '( dport = :443 )' | wc -l
28229
curl -v https://api.lab/
*   Trying 10.0.3.20:443...
* Immediate connect fail for 10.0.3.20: Cannot assign requested address
curl: (7) Failed to connect to api.lab port 443 after 0 ms: Couldn't connect to server

Cannot assign requested address is EADDRNOTAVAIL - the kernel's name for this error (an errno: the numbered error code a system call returns): no free source port for that destination. Java says it as java.net.NoRouteToHostException: Cannot assign requested address, which sends people looking at routing. Other destinations still work fine, which is the second clue.

The fix is connection reuse, not sysctls

HTTP keep-alive means "keep this connection open after the response, I will send another request on it". A connection pool is a small set of open connections an app keeps and lends to each request in turn. Together they send thousands of requests over a handful of connections. No new tuples, no TIME-WAIT pile, and no TCP + TLS handshake per request (which is also most of the latency for small requests). The sysctls people reach for:

net.ipv4.ip_local_port_range   wider range: more headroom, same bug
net.ipv4.tcp_tw_reuse          lets NEW outgoing connections reuse a TIME-WAIT
                               tuple when TCP timestamps make it safe. Default 2
                               = loopback only. Setting 1 is reasonable on
                               clients; it is a mitigation.
net.ipv4.tcp_tw_recycle        removed from Linux in 4.12. It broke clients
                               behind NAT. If a blog post tells you to set it,
                               close the blog post.

Behind a NAT gateway (Chapter 8), the same arithmetic happens on the NAT's ports. SNAT (source NAT) rewrites your private source address and port into the gateway's public ones, so every machine behind it shares one pool of ports: SNAT port exhaustion. Cloud NAT gateways often hand each machine only a small fixed share. A fleet of apps that does not reuse connections exhausts it and outbound calls start failing now and then. Same fix: reuse connections. Then more public addresses on the gateway.

The server side: backlog and the accept queue

A listening socket has a queue of completed connections waiting for the application to accept() them (9.5). Its size is the backlog the application asked for, capped by the kernel setting net.core.somaxconn:

$ sysctl net.core.somaxconn
net.core.somaxconn = 4096
$ ss -ltn
State  Recv-Q Send-Q Local Address:Port  Peer Address:Port Process
LISTEN 0      511          0.0.0.0:80         0.0.0.0:*
LISTEN 0      100          0.0.0.0:8090       0.0.0.0:*
LISTEN 0      4096               *:22               *:*

For a LISTEN row, Send-Q is the backlog. nginx asks for 511; the Java web server inside the app on 8090 for 100 (its accept-count setting); and port 22 shows systemd's default for a socket unit, 4096 (sshd is socket-activated on Ubuntu - systemd listens and starts sshd on demand; started by hand it asks for 128).

A Java web app handles requests with a fixed set of worker threads (a thread is a line of work inside one process; each worker handles one request at a time). When every worker thread is busy (say, all blocked waiting on a database), nobody calls accept(). The queue fills:

$ ss -ltn '( sport = :8090 )'
State  Recv-Q Send-Q Local Address:Port  Peer Address:Port Process
LISTEN 101    100          0.0.0.0:8090       0.0.0.0:*

Recv-Q 101 on a queue of 100 - full. From then on the kernel drops new SYNs (or the final ACK) instead of refusing them, so clients see a timeout, not "refused", against a port that is visibly listening. The counters prove it:

$ nstat -az TcpExtListenOverflows TcpExtListenDrops
#kernel
TcpExtListenOverflows           1843               0.0
TcpExtListenDrops               1843               0.0

nstat prints the kernel's network counters: -a absolute totals (not the change since last run), -z include counters that are zero. Without -a it prints the change since you last ran it: run it twice a few seconds apart and see if it is still climbing. Raising the backlog only buys queue; the fix is whatever is holding the workers - usually a missing timeout on a downstream call (a call to a service further down the chain).

Keepalive, and the load balancer's idle timeout

TCP keepalive (not the same thing as HTTP keep-alive) sends an empty probe on an idle connection so both ends (and everything in between) know it is alive. Linux defaults are useless for this:

$ sysctl net.ipv4.tcp_keepalive_time net.ipv4.tcp_keepalive_intvl net.ipv4.tcp_keepalive_probes
net.ipv4.tcp_keepalive_time = 7200
net.ipv4.tcp_keepalive_intvl = 75
net.ipv4.tcp_keepalive_probes = 9

First probe after two hours of idle (7200 s), then every 75 s, 9 times - and only on sockets whose application enabled keepalive at all (the SO_KEEPALIVE socket option).

Load balancers, NAT gateways and firewalls remember every connection that passes through them in a table, and forget idle ones after an idle timeout - 4 minutes is a common cloud default. They do it silently. Neither end is told. The next time either side writes, it gets an RST - "Connection reset by peer" out of nowhere, on the first request after a quiet period. Database pools, message consumers, long streams: anything long-lived and quiet.

Fixes: keepalives below the timeout, set in the client or pool (for example a 2-minute keepalive, or a pool that checks or throws away idle connections); or raise the LB's idle timeout; or have the LB send an RST when it forgets a flow, so clients at least learn about it immediately. sudo sysctl -w net.ipv4.tcp_keepalive_time=120 (-w = write) changes the default for applications that enable keepalive; put it in a file under /etc/sysctl.d/ to keep it after a reboot.

ss -o shows it working:

ESTAB 0 0 10.64.0.2:45120 10.0.3.20:443 timer:(keepalive,1min48sec,0)

Later (Ch 22): Azure's load balancer is one of these devices, with a 4-minute default idle timeout and its own SNAT port limits - the arithmetic above is exactly what you will size there.

What you can now do

Why it helps

This is behind some of the nastiest intermittent incidents. A service calling a partner API through a NAT gateway starts failing a few percent of calls at peak: SNAT port exhaustion, caused by a client that creates a new HTTP client per request. Java throws NoRouteToHostException: Cannot assign requested address and people spend hours on routing tables. On the server side, the reports service stops answering new users while ss shows the port listening: its accept queue is full because every worker thread waits on a database with no timeout.

Knowing the arithmetic lets you reject "just widen the port range" in a PR review and ask for connection pooling and downstream timeouts instead.

Commands in this lesson

sysctl ss nstat

FAQ

Aren't there 65,535 ports? Why only about 28,000?

Source ports for outgoing connections come from net.ipv4.ip_local_port_range, which defaults to 32768-60999 on Linux: 28,232 ports. And the limit is per destination IP and port, because a connection is identified by the 4-tuple (source IP, source port, destination IP, destination port). Connections to your database and to an API use separate pools of tuples, so they do not compete with each other.

Should I set tcp_tw_reuse or tcp_tw_recycle?

Never tcp_tw_recycle: it broke clients behind NAT and was removed from Linux in 4.12, so any guide recommending it is obsolete. tcp_tw_reuse defaults to 2 (loopback only); setting 1 lets new outgoing connections reuse a TIME-WAIT tuple when timestamps make it safe. It is a reasonable mitigation on a client, but it hides the real bug: not reusing connections.

What is SNAT port exhaustion?

When machines with private addresses reach the internet through a NAT gateway, their private IP and port are translated to the gateway's public IP and one of its ports (source NAT). Cloud gateways often hand each machine a fixed, small number of those ports. Every outbound connection to the same destination consumes one. A fleet that does not reuse connections burns through them, and outbound calls start failing now and then. Fix connection reuse first; then give the gateway more public IPs.

Why do clients see timeouts instead of "refused" when the server is overloaded?

Because when the accept queue is full, the Linux kernel drops new SYNs (or the final ACK) instead of rejecting them. The port is visibly listening, so nobody suspects it. You spot it with ss -ltn showing Recv-Q above Send-Q on the listening socket, and with nstat -az TcpExtListenOverflows TcpExtListenDrops climbing between two runs. Raising the backlog only buys a bigger queue.

Does TCP keepalive keep connections alive through load balancers by default?

No. On Linux the first keepalive probe goes out after 7200 seconds of idle (tcp_keepalive_time), and only on sockets whose application enabled SO_KEEPALIVE. Cloud load balancers commonly forget idle flows after 4 minutes. So keepalive must be configured below that timeout, ideally in the client or connection pool, or the pool must validate and evict idle connections.

In an interview Junior

A client suddenly fails with "Cannot assign requested address" to one server, while everything else works. What is going on?

Ephemeral port exhaustion. A connection is a 4-tuple (source IP, source port, destination IP, destination port); towards one destination only the source port varies, from net.ipv4.ip_local_port_range (32768-60999, about 28,000 ports).

A client that opens a new connection per request and closes it first leaves each one in TIME-WAIT for 60 seconds, holding its port: about 470 new connections a second to one destination and connect() fails with EADDRNOTAVAIL. Other destinations still work - the second clue.

Prove it: ss -Htan state time-wait '( dport = :443 )' | wc -l close to the range size.

Fix: connection reuse - HTTP keep-alive and a connection pool. A wider port range or tcp_tw_reuse only buys headroom. Behind a NAT gateway the same arithmetic hits its SNAT ports.

Also asked: Why does the first request after a quiet period sometimes fail with "connection reset"? · What is the listen backlog, and what happens when the accept queue is full? · What is the difference between TCP keepalive and HTTP keep-alive?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.