Why this matters
"The service runs out of connections after six hours" and "the box has 20,000 connections open, is that bad?" are both answered by one command, ss, if you can read what it prints. Every connection is in a state, and a pile-up in one particular state tells you whose bug it is.
What you need to know already: the handshake and the SYN/ACK/FIN/RST flags (9.1), sudo ss -tlnp (9.1), and file descriptors and their limits (3.16) - every socket a process holds is one of its file descriptors.
The states you will actually see
TCP keeps a small state machine for every connection: a fixed list of states and the events that move it from one to the next. ss prints the current state.
LISTEN bound and accepting new connections
SYN-SENT we sent a SYN, no answer yet (a stuck connect())
SYN-RECV we got a SYN, sent SYN-ACK, waiting for the final ACK
ESTAB connected
FIN-WAIT-1 we closed, waiting for the peer to ACK
FIN-WAIT-2 our close was ACKed, waiting for the peer's FIN
TIME-WAIT we closed FIRST and it is done; kept ~60s to absorb stray packets
CLOSE-WAIT the PEER closed; waiting for OUR application to call close()
LAST-ACK we closed after the peer; waiting for the final ACK
Who ends up where depends on who closes first:
active closer (sends FIN first) FIN-WAIT-1 -> FIN-WAIT-2 -> TIME-WAIT -> gone
passive closer (receives FIN) CLOSE-WAIT -> (app calls close) -> LAST-ACK -> gone
Two rules fall out of that:
- TIME-WAIT belongs to whoever closed first, and lasts 60 seconds on Linux (a fixed constant,
TCP_TIMEWAIT_LEN;tcp_fin_timeoutdoes not change it - that is the FIN-WAIT-2 timeout). A busy client that closes its connections has thousands. That is normal. - CLOSE-WAIT waits for your code. The kernel has done its part; only the process holding the file descriptor can finish it, by calling
close()(the system call - a request from a program to the kernel, the things strace showed you in 3.12 - that closes a file descriptor). A CLOSE-WAIT that does not go away is a socket leak (sockets opened and never closed) - always the application's bug.
SYN-SENT piling up is the third one to recognise: the application is stuck in connect() to something that does not answer - a DROP, the timeout row of the previous lesson, seen from the inside.
Reading ss
$ ss -tan
State Recv-Q Send-Q Local Address:Port Peer Address:Port Process
LISTEN 0 4096 0.0.0.0:8080 0.0.0.0:*
LISTEN 0 4096 *:22 *:*
ESTAB 0 0 10.64.0.2:22 10.64.0.1:51234
ESTAB 0 0 10.64.0.2:45120 10.0.3.20:443
TIME-WAIT 0 0 10.64.0.2:40117 93.184.216.34:443
CLOSE-WAIT 1 0 10.64.0.2:8080 10.0.2.14:40003
The flags: -t TCP, -u UDP, -a all states (without it, listening sockets and TIME-WAIT are hidden), -l listening only, -n numeric ports, -p the owning process (root needed for other users' processes), -o timers, -e extended (uid, inode), -H no header, -s summary. They combine: -tan is -t -a -n.
Each row is one socket: its state, two queue counters, our address and port (Local), the other side's (Peer). Row 3 is your own SSH session from the Mac; row 5 is this box calling a service on port 443.
The two queue columns (a queue is a waiting line of bytes or connections) mean different things depending on the state:
ESTAB Recv-Q = bytes received that the app has not read yet
Send-Q = bytes sent that the peer has not acknowledged
LISTEN Recv-Q = connections waiting in the accept queue RIGHT NOW
Send-Q = the size of that queue (the backlog)
The accept queue: the kernel finishes handshakes on its own and parks the new connections in a line; the application takes the next one with the accept() system call. The backlog is how long that line may get.
- A growing Recv-Q on ESTAB - the application is not reading (busy, blocked).
- A growing Send-Q on ESTAB - the peer or the network is not keeping up.
- LISTEN with Recv-Q at Send-Q + 1 - the accept queue is full: the application has stopped calling
accept(). New connections are dropped. The ports lesson has the incident. CLOSE-WAIT 1 0- one byte unread: typically the peer's FIN, still sitting there because nobody read (and closed) the socket.
Filters, instead of grep
ss can filter by itself. state NAME keeps one state; dst / src match the peer's / our address; dport / sport match the destination / source port. Inside the quotes and parentheses you write a small expression, like the find expressions of 4.15. Letting ss filter beats grep, because grepping for 5432 also matches IPs and other columns.
ss -tan state established one state
ss -tan state close-wait
ss -tan state time-wait
ss -tan state connected everything except LISTEN and closed
ss -tan exclude listening
ss -tn dst 10.0.3.12 to one host
ss -tn '( dport = :5432 )' to a port
ss -tn '( sport = :8080 )' from our port 8080 (inbound to us)
ss -tn state established '( dport = :443 or dport = :5432 )'
With a single state filter, ss drops the State column - it would be the same on every line:
# during the CLOSE-WAIT leak in the mission below
ss -tan state close-wait | head -3
Recv-Q Send-Q Local Address:Port Peer Address:Port Process
1 0 10.64.0.2:8080 10.0.2.14:40003
1 0 10.64.0.2:8080 10.0.2.15:40004
And the header line is still there, so | wc -l counts one too many. Use -H:
# during the leak
ss -Htan state close-wait | wc -l
340
Who owns it
$ sudo ss -tlnp
State Recv-Q Send-Q Local Address:Port Peer Address:Port Process
LISTEN 0 100 0.0.0.0:8080 0.0.0.0:* users:(("java",pid=1210,fd=41))
LISTEN 0 511 0.0.0.0:80 0.0.0.0:* users:(("nginx",pid=905,fd=6))
users:(("java",pid=1210,fd=41)) - process name, PID, and the file descriptor number, which you can find again in /proc/1210/fd/41. Without sudo, the Process column is empty for sockets owned by other users - the same trap as lsof in 3.12. (java here is the orders service - a Java program, run by the JVM you met in 5.13.)
Timers
$ ss -tano state time-wait | head -3
Recv-Q Send-Q Local Address:Port Peer Address:Port Process
0 0 10.64.0.2:40117 93.184.216.34:443 timer:(timewait,43sec,0)
0 0 10.64.0.2:40118 93.184.216.34:443 timer:(timewait,51sec,0)
timer:(name,time left,retries): timewait,43sec - 43 seconds until it is gone. Other timers: keepalive (when the next keepalive probe goes out - 9.8), on (a retransmission timer running - something is not being ACKed), persist (the peer said its window is full; we are asking again).
The summary
$ ss -s
Total: 541
TCP: 355 (estab 3, closed 40, orphaned 0, timewait 40)
Transport Total IP IPv6
RAW 1 0 1
UDP 4 3 1
TCP 315 312 3
INET 321 315 6
FRAG 0 0 0
closed includes TIME-WAIT. orphaned counts sockets no process holds any more (closed by the app, still finishing in the kernel). The table below splits the totals by protocol (RAW, UDP, TCP) and by IPv4/IPv6. The first command in any "we are out of connections" incident: it tells you which state to dig into.
What you can now do
- Read an
ss -tanrow: state, queues, local and peer address. - Say whose problem a pile-up is: TIME-WAIT (usually fine), CLOSE-WAIT (the app's bug), SYN-SENT (something drops our connects).
- Count one state towards one port with
ss -Hfilters andwc -l.