OnCallReady

Lesson 9.1 · TCP, TLS & HTTP · 24 min read

The handshake, and the three ways a connection fails

In plain words

Imagine phoning a shop. You dial, it rings, someone picks up and says "hello", you say "hello" back, and only then do you start talking. That is the TCP handshake: SYN, SYN-ACK, ACK.

Now the three ways a call goes wrong. A recorded voice says "this number is not in service" straight away: someone answered, there is just nobody at that extension. That is refused. The phone rings and rings and nobody picks up: timed out. You were talking and the line suddenly goes dead: reset.

On oncall-lab, nc -zv and curl show the same three outcomes, and the time they take is the clue: instant means a host answered, slow means nothing did.

Why this matters

You run curl against a service and it fails. The error is one of three kinds, and each kind points at a different culprit: the service itself, a firewall, or something in the middle. This lesson teaches you to tell them apart in two seconds - from the message, and from how long it took to arrive.

What you need to know already: IP addresses, ports and routes (Chapter 8), the nc and curl probes from the "I cannot reach X" method (8.29), and that 127.0.0.1 is the box talking to itself (loopback).

The words first

The handshake, as it looks on the wire

client                                  server
  SYN            seq=x          ->
                               <-      SYN-ACK   seq=y, ack=x+1
  ACK            ack=y+1        ->
  ... data ...
  FIN            ->                               (client closes first)
                               <-      ACK
                               <-      FIN
  ACK            ->

SYN, SYN-ACK, ACK: that is the three-way handshake, and it costs one RTT. Closing is separate: each side sends a FIN and the other ACKs it.

tcpdump (you used it in 8.28 to count DNS queries; 9.11 covers it fully) prints one line per packet. This is what it shows for nc -z github.com 443 (nc -z: open a connection to that port and close it again, sending no data):

IP 10.64.0.2.41250 > 140.82.121.3.443: Flags [S], seq 1829301, win 64240, options [mss 1460,sackOK,TS val 1301 ecr 0,nop,wscale 7], length 0
IP 140.82.121.3.443 > 10.64.0.2.41250: Flags [S.], seq 2910283, ack 1829302, win 65160, options [mss 1460,sackOK,TS val 2283 ecr 1301,nop,wscale 7], length 0
IP 10.64.0.2.41250 > 140.82.121.3.443: Flags [.], ack 1, win 502, options [nop,nop,TS val 1302 ecr 2283], length 0

Reading one line, left to right: source-ip.port > destination-ip.port, the flags, the sequence and ack numbers, win (the window: how many bytes the sender can accept right now), options, and length (bytes of data - 0 in a handshake). Our side uses port 41250: the kernel picked a random free source port for us; 443 is the service's port.

You do not need to recite this. You need the consequences: the three ways it goes wrong, which tell you where to look.

The most useful table in networking

"Listening" means a server program has asked the kernel to accept connections on a port. A load balancer (LB) is a machine in front of several servers that spreads incoming connections across them.

SymptomWhat happened on the wireUsual cause
Connection refused, instantlySYN out, RST backnothing listening on that port; service down; listening on 127.0.0.1 only
Connection timed out, slowlySYN out, nothing back, SYN retrieda firewall silently dropping it; wrong route; host gone
Connection reset by peer, mid-flightan RST after the connection workedan LB forgot an idle connection, the server crashed or was killed, a proxy closed it

The timing is the diagnosis. Instant means something answered - there is a route, the host is up, the port is closed. Slow means nothing answered at all.

Refused, on every tool

The two probes you will use all chapter:

$ nc -zv oncall-lab 9999
nc: connect to oncall-lab (127.0.1.1) port 9999 (tcp) failed: Connection refused

$ curl http://oncall-lab:9999/
curl: (7) Failed to connect to oncall-lab port 9999 after 0 ms: Couldn't connect to server

$ curl -v http://oncall-lab:9999/
*   Trying 127.0.1.1:9999...
* connect to 127.0.1.1 port 9999 from 127.0.0.1 port 51712 failed: Connection refused
* Failed to connect to oncall-lab port 9999 after 0 ms: Couldn't connect to server

And on the wire:

IP 127.0.0.1.51712 > 127.0.1.1.9999: Flags [S], seq 3318821, ...
IP 127.0.1.1.9999 > 127.0.0.1.51712: Flags [R.], seq 0, ack 3318822, win 0, length 0

[R.] - reset. The kernel on the other side said "nothing here" in one round trip. after 0 ms is the tell in curl's message.

Timed out, on every tool

$ nc -zv -w 3 db.lab 5432
nc: connect to db.lab (10.0.3.12) port 5432 (tcp) timed out: Operation now in progress

$ curl --connect-timeout 3 http://db.lab:5432/
curl: (28) Failed to connect to db.lab port 5432 after 3002 ms: Timeout was reached

(--connect-timeout 3: curl gives up connecting after 3 seconds.)

Without a timeout of your own, the kernel decides - and it is patient. sysctl NAME prints a kernel setting (you changed vm.swappiness the same way in Chapter 5):

$ nc -zv db.lab 5432
nc: connect to db.lab (10.0.3.12) port 5432 (tcp) failed: Connection timed out
$ sysctl net.ipv4.tcp_syn_retries
net.ipv4.tcp_syn_retries = 6

Six retries with exponential backoff (each wait twice as long as the last): SYNs at 0, 1, 3, 7, 15, 31 and 63 seconds, then give up at about 127 seconds. That is the "it hangs for two minutes" everyone has seen. On the wire it is the same SYN, again and again:

IP 10.64.0.2.40412 > 10.0.3.12.5432: Flags [S], seq 771820, ...
IP 10.64.0.2.40412 > 10.0.3.12.5432: Flags [S], seq 771820, ...
IP 10.64.0.2.40412 > 10.0.3.12.5432: Flags [S], seq 771820, ...

Same source port, same seq: retransmissions (the same packet sent again), not new attempts.

Always set a connect timeout in real clients (--connect-timeout for curl, and the connect-timeout setting every database or HTTP library has). A default of "the kernel's 127 seconds" turns one dead dependency (a service your app calls) into an app full of stuck requests.

REJECT vs DROP

A firewall is a set of rules, on a host or on a box in the path, that decides which packets may pass. For a packet it does not allow, a rule either answers for the host or stays silent:

REJECT (with tcp-reset)  -> the client sees "refused", instantly
REJECT (icmp)            -> "refused" or "No route to host", instantly
DROP                     -> the client sees a timeout, slowly

Most firewalls DROP, and the cloud ones always do. So "it just hangs" is the signature of a firewall problem, and "refused" almost never is. Two seconds of observation saves you from debugging the wrong layer.

Reset by peer

The connection worked, then an RST arrived (peer = the other end):

# an export service behind an LB that dropped the idle connection (not on this box)
curl http://10.0.3.77:8080/export
curl: (56) Recv failure: Connection reset by peer

Causes, most common first: a load balancer or NAT (Chapter 8) that forgot the connection after it sat idle (the idle timeout, next lesson), the server process crashed or was killed while you were talking, a proxy (a middleman that forwards traffic, 9.23) closed it, or encrypted TLS (9.15) spoken to a port that expects plain TCP (or the reverse).

"Works on localhost only": the bind address

A server program chooses which address it listens on - it binds to it. ss (socket statistics; you glimpsed it in Chapter 3) shows every listening program. -t TCP only, -l listening sockets only, -n numbers instead of names, -p the owning process (needs sudo for other users' processes). A socket is the kernel's object for one end of a connection, or for a listener.

$ sudo ss -tlnp
State  Recv-Q Send-Q Local Address:Port  Peer Address:Port Process
LISTEN 0      4096       127.0.0.1:9100       0.0.0.0:*    users:(("prometheus-node",pid=2210,fd=3))
LISTEN 0      4096               *:22               *:*    users:(("sshd",pid=700,fd=3),("systemd",pid=1,fd=58))

The columns: State, two queue counters (next lesson), Local Address:Port (where it listens), Peer Address:Port (* = anyone may connect) and Process (name, PID, file descriptor). The first line is a metrics exporter: a small program that publishes the box's numbers (load, memory) over HTTP so a monitoring server can collect them. You fix this exact one in 9.4.

sshd listens on * - every address, IPv4 and IPv6 (0.0.0.0 is the IPv4-only spelling). The exporter only on 127.0.0.1. So:

$ curl -s http://localhost:9100/metrics | head -2
# HELP node_load1 1m load average.
# TYPE node_load1 gauge
$ curl http://10.64.0.2:9100/metrics
curl: (7) Failed to connect to 10.64.0.2 port 9100 after 0 ms: Couldn't connect to server

Running, healthy, reachable from the box, refused from everywhere else - including from the box's own LAN address. This is the number one cause of "it works on my machine and nowhere else". The fix is always the bind address (0.0.0.0, :: or an empty host like :9100), never the firewall.

*:9100 in the Local Address column means "all addresses, IPv4 and IPv6" - what many programs show when they are told to listen on :9100.

Later (Ch 10): inside a container, 127.0.0.1 means the container itself, so an app bound there cannot be reached even from its own host - the same bug, one level deeper.

What you can now do

Why it helps

This is the first fork in almost every connectivity page. The monitoring server says a service is down while it runs fine: ss -tlnp shows it bound to 127.0.0.1, so every connection to the box's real IP is refused, and you fix the bind address, not the firewall. An app call to a database hangs for two minutes: that is SYNs being dropped by a firewall, so you stop reading application logs and go look at the firewall rules. A batch export dies with "reset by peer" after running fine for a while: think load balancer idle timeout or a crashed upstream.

Knowing these three signatures means you pick the right team and the right layer within seconds, instead of restarting things and hoping.

Commands in this lesson

nc curl sysctl ss

FAQ

Why does a dropped connection take about two minutes to fail?

Because the kernel retries the SYN on its own, with exponential backoff. With net.ipv4.tcp_syn_retries = 6 it sends at 0, 1, 3, 7, 15, 31 and 63 seconds and gives up at roughly 127 seconds. The application just sits in connect() the whole time. That is why every real client needs its own connect timeout: --connect-timeout in curl, and the connect-timeout setting of every database and HTTP library an app uses.

If the firewall blocks me, why don't I get "refused"?

Because most firewalls DROP rather than REJECT. A REJECT sends back an RST or an ICMP error, so you would see refused or "No route to host" immediately. A DROP sends nothing, so you get a timeout. Cloud firewalls always drop. So "it hangs" usually points at a firewall, and "refused" almost never does.

The app works with curl on localhost but not from another machine. Is it the firewall?

Usually not. Check the bind address first with sudo ss -tlnp. If the Local Address is 127.0.0.1:9100, the process only accepts connections on loopback, and connections to the box's real IP are refused instantly, even from the box itself. The fix is to bind to 0.0.0.0, :: or an empty host like :9100. A firewall problem would look like a timeout, not a refusal.

What does the dot in [S.] mean in tcpdump?

The dot is the ACK flag. [S] is a bare SYN (the client opening), [S.] is SYN plus ACK (the server accepting and acknowledging the client's SYN), [.] is an ACK alone, and [R.] is a reset with ACK, the usual "port closed" reply. The server's ack number is the client's sequence number plus one, which means "I received your SYN".

What does "Connection reset by peer" actually tell me?

That the connection was established and then the other side, or something pretending to be it, sent an RST. The usual causes are a load balancer or NAT that dropped the idle flow (4 minutes is a common cloud default), a server process that crashed or was killed mid-request, a proxy that closed the connection, or a protocol mismatch like speaking TLS to a plain TCP port. It is never a routing or DNS problem, because the connection already worked.

In an interview Junior

Explain the TCP three-way handshake, and what "connection refused" and "connection timed out" tell you.

Before data flows, the kernels set up the connection: the client sends SYN ("let's talk", its starting seq), the server answers SYN-ACK (its own seq, ack = client seq + 1), the client sends ACK. It costs one RTT.

The failures, from the wire:

The timing is the diagnosis: instant means something answered.

Also asked: What does it mean when a service listens on 127.0.0.1 instead of 0.0.0.0? · What is the difference between a firewall that rejects and one that drops? · How do you test whether a TCP port is reachable from a box?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.