OnCallReady

Chapter 9 TCP, TLS & HTTP

The three connection failures and how to tell them apart, socket states that are your bug, packets on the wire, MTU black holes, certificate chains, proxies, and finding where latency actually lives.

In plain words

Think of sending a letter to a company. First the post office has to agree there is a road to the building and someone at the door (that is TCP: a phone call where both sides say "hello, can you hear me?" before talking). Then you check the badge of the person who opens the door, and that badge was stamped by an office you already trust (that is TLS: proving who is on the other end). Only then do you hand over the actual letter in a fixed format, "please send me page /cart" (that is HTTP).

On oncall-lab each layer has its own tool: ss and tcpdump for TCP, openssl s_client and openssl x509 for TLS, curl -v and curl -w for HTTP. When something is broken, it is broken at exactly one of these layers, and the error message tells you which one.

Why it matters on call

Most "the service is down" pages are not about the service. They are a firewall that drops SYNs (a timeout), an app bound to 127.0.0.1 (refused), a certificate that expired at midnight, an intermediate nobody deployed, an nginx 504 because the upstream waits on a database, or a client that opens a new connection per request and runs out of ports.

This chapter gives you the reflex to read the symptom and name the layer in seconds: instant vs slow failure, curl exit 7 vs 28 vs 60, 502 vs 504 vs 499, CLOSE-WAIT vs TIME-WAIT. That is what the on-call engineer is paid for, it is what you need to review a proxy or load balancer change, and "what happens when you type a URL" is the most asked networking interview question there is. It comes right after addressing and DNS because every one of these steps starts with a name becoming an IP.

Lessons

  1. The handshake, and the three ways a connection fails
  2. Socket states, and reading ss properly
  3. Ports, queues and exhaustion
  4. Reading tcpdump
  5. MTU, MSS and black holes
  6. TLS: the handshake, the chain, and trust
  7. Certificates with openssl, and the JVM truststore
  8. HTTP on the wire, with curl -v
  9. Reverse proxies, load balancers and forward proxies
  10. Where latency lives: curl -w and the road to the first byte

23 hands-on labs (missions, incidents and drills) run in the terminal: Open this chapter in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.

Questions people ask

Why learn TCP and TLS by hand when cloud services and platforms hide them?

They hide the configuration, not the failures. A service in the cloud that cannot reach a database still sends a SYN that nobody answers, a managed proxy still presents a certificate chain, and a cloud load balancer still silently forgets idle connections after a few minutes. The platforms you will meet later give new names to the same old things: a cloud firewall rule is a DROP rule, a platform "health probe" is an active health check, a managed gateway is a reverse proxy. When the dashboard is green and users see errors, only the wire-level view tells you what is really happening.

Which tool do I reach for first?

curl -v against the failing URL, from the box that is failing. It narrates DNS, the TCP connect, the TLS handshake and the HTTP exchange in order, and it stops exactly at the broken step. Then go deeper at that layer: ss for socket states and queues, tcpdump when you need to prove what was or was not on the wire, openssl s_client for the certificate chain, curl -w for timings. Starting with tcpdump is usually too low; starting with application logs is usually too high.

What is the difference between "refused", "timed out" and "reset"?

Refused comes back instantly: your SYN reached a host whose kernel answered with an RST because nothing listens on that port (or it listens only on 127.0.0.1). Timed out is slow: nothing answered at all, so a firewall dropped the packet, a route is wrong or the host is gone. Reset arrives after the connection already worked: a load balancer forgot the flow, the process died, or a proxy closed it. The timing alone tells you which layer to debug.

Is TLS the same thing as HTTPS?

HTTPS is HTTP carried inside a TLS connection on port 443. TLS itself works under any protocol: PostgreSQL, LDAP, SMTP (mail) and message queues all use it too, so everything you learn about chains, SANs, expiry and trust stores applies to database connections as much as to web traffic. The Java error "PKIX path building failed" is the same problem whether the JVM was calling a web API or a database.

Do I need to memorise the TCP state machine and all the flags?

No. You need a handful of consequences: TIME-WAIT belongs to whoever closed first and is normal; CLOSE-WAIT that grows is always an application bug; SYN-SENT piling up means something drops your SYNs; a LISTEN socket with Recv-Q above its backlog means the app stopped accepting. In tcpdump, recognise [S], [S.], [.], [P.], [F.] and [R]. That covers nearly every real incident.