OnCallReady

Lesson 8.22 · Addressing & DNS · 13 min read

Records, TTLs, negative caching and split horizon

In plain words

A phone book can hold different kinds of entries. "Maria: 555-1234" is a direct number (an A record). "The bakery: see Maria" is a redirect to another entry (a CNAME). "Post for this building goes to the concierge" is an MX. And every copy of the phone book has a "valid until" date, so people who photocopied an old page keep using it until then, even after you change your number. That is the TTL.

People also remember "no such person" for a while, which is negative caching. And some buildings keep two phone books: one for residents, one for visitors, with different answers for the same name. That is split horizon, which you saw when 8.8.8.8 had never heard of api.lab.

Why this matters

A service moves to a new server, DNS is changed, and half the users keep hitting the old one for an hour. Or you create a new name and it "does not exist" for five minutes. Both follow directly from how records and TTLs work - and both are avoidable if you plan.

What you need to know already: 8.16 (zones, authoritative servers, TTL) and 8.18 (reading dig).

The record types

A zone holds records: one fact each (a name, a type, a TTL, the data):

A       name -> IPv4 address
AAAA    name -> IPv6 address
CNAME   name -> another name (an alias). The resolver restarts the lookup.
NS      which servers are authoritative for a zone (and for delegations)
SOA     one per zone: primary server, admin, serial, timers, negative TTL
MX      mail servers with a priority (LOWER number = preferred)
TXT     free text: used to prove you own a domain, and for mail settings
SRV     where a service runs: priority weight port target
PTR     IP -> name, under in-addr.arpa (reverse DNS)
CAA     which companies may issue security certificates for the name

An SRV record, read field by field:

_postgresql._tcp.db.lab.  300  IN  SRV  0 5 5432 db.lab.
                                        | |  |    target host
                                        | |  port
                                        | weight (among same priority)
                                        priority (lower first)

The name says which service (_postgresql, a database) over which transport (_tcp), in which domain.

CNAME has rules, and they bite

  1. A CNAME cannot sit next to any other record at the same name. If www is a CNAME, it cannot also have TXT or MX records.
  2. No CNAME at the zone apex. The apex is the zone's own name, example.com itself. It must have SOA and NS records, so by rule 1 it cannot be a CNAME. You cannot point example.com at a hosting company's name with a CNAME.

That is why DNS providers invented their own workarounds ("alias" or "flattened" records): their server looks up the target itself and hands out plain A records.

A CNAME chain costs a lookup per hop and every hop has its own TTL. The answer's lifetime is the shortest TTL in the chain.

The SOA record

lab.  3600  IN  SOA  ns1.lab. hostmaster.lab. 2026092301 7200 3600 1209600 300
                     |        |               |          |    |    |       |
                     primary  admin (hostmaster@lab)     |    |    |       minimum
                                              serial     |    |    expire
                                                         refresh retry

TTL: why your change did not take effect

Every resolver that holds the old answer keeps it until its TTL runs out. You cannot flush someone else's cache. So a change to a record with TTL 3600 takes up to an hour to reach everyone, and some programs (Java services, nginx, 8.16) hold it longer.

The migration procedure:

T-48h   check the current TTL:  dig +noall +answer @ns1.lab orders.lab
T-24h   lower it to 60 (or 300)
        lowering a TTL is itself subject to the OLD TTL - resolvers that cached
        the record keep the 3600 until it expires. So do it a full old-TTL
        (plus margin) before the change, not an hour before.
T-0     change the record
T+5m    verify from several resolvers; watch the old server's logs
T+1d    raise the TTL back (60s forever = 60x the query load, for nothing)

(T-24h = 24 hours before the change, T+5m = 5 minutes after.)

Negative caching: NXDOMAIN sticks too

When a name does not exist, resolvers remember that as well - negative caching - for:

negative TTL = min(SOA record's own TTL, SOA minimum field)

For lab. that is min(3600, 300) = 300 seconds. The classic trap:

  1. The app starts and looks up newsvc.lab - it does not exist yet: NXDOMAIN.
  2. You create the record a minute later.
  3. For up to five more minutes, that app's resolver keeps answering NXDOMAIN.

"I created the record and it still says it does not exist" is almost always this. On your own box, sudo resolvectl flush-caches fixes it; on everyone else's, you wait.

Split-horizon DNS

The same name gives different answers depending on who asks - split horizon. Company resolvers serve an internal view; the internet sees a public one, or nothing:

$ dig +short api.lab                 # from inside: the internal resolver
10.0.3.20
$ dig +short @8.8.8.8 api.lab        # from the internet's point of view
$ dig @8.8.8.8 api.lab | grep status
;; ->>HEADER<<- opcode: QUERY, status: NXDOMAIN, id: 9912

(8.8.8.8 is Google's public resolver - it only knows the public internet.) It is deliberate, and common:

The classic failure: a machine is set up to use a public resolver (or any resolver that does not have the internal view). It gets the public address, connects there, and is refused, because the service only accepts connections on the private one. The network looks fine; DNS gave the wrong view.

When you debug DNS across networks, always ask: which resolver did this client use, and which view does that resolver have?

Later (Ch 23): cloud "private endpoints" work exactly like this, and the wrong-resolver failure is one of the most common cloud tickets.

What you can now do

Why it helps

Record rules turn into incidents regularly. A migration where the TTL was lowered an hour before the change instead of a day before, so half the users hit the old server for hours. A new service that "does not exist" for five minutes after you created it, because a client asked too early and cached NXDOMAIN. A team trying to CNAME their bare domain to a hosting company's name, which DNS forbids.

Split horizon is the one that costs the most time: a machine using a resolver without the internal view gets a public address, connects to the wrong place and is refused. The network looks fine; DNS gave the wrong view. Knowing this turns a day of firewall debugging into one dig against the right server.

Commands in this lesson

dig

FAQ

Why can't I put a CNAME on example.com itself?

Because a CNAME cannot sit next to any other record at the same name, and the zone apex must have SOA and NS records. So the apex can never be a CNAME. DNS providers work around it on their side with "alias" or "flattened" records: their server looks up the target itself and returns plain A records, which looks normal to every client.

I lowered the TTL right before the migration. Why did it not help?

Lowering a TTL is itself subject to the old TTL. Resolvers that cached the record with TTL 3600 keep it, including the old TTL, for up to an hour. They only learn the new, short TTL when their cached copy expires. So lower it at least one full old-TTL (plus margin) before the change, typically a day ahead, then raise it back afterwards.

How long is an NXDOMAIN cached?

For the negative TTL, which is the smaller of the SOA record's own TTL and the SOA's minimum field (the last number). With lab. 3600 SOA ... 300 that is 300 seconds. So if a client asks for a name before you create it, it keeps hearing "does not exist" for up to five minutes afterwards. Flush the caches you control; the others you wait out.

What is the SOA serial for, and what happens if I forget to bump it?

The serial is how backup (secondary) servers know the zone changed. They check the primary every refresh interval and copy the zone only if the serial went up. Forget to bump it and the secondaries keep serving the old data, so some queries get the new answer and some the old, depending on which server answers. The convention is YYYYMMDDnn. Hosted DNS services bump it for you.

What does "which view does that resolver have" mean in practice?

With split horizon, the answer depends on who asks. Only resolvers that have the internal view (your company's, or the lab's upstream at 10.64.0.1) know internal names or return internal addresses. A machine set up with 8.8.8.8, or any resolver without that view, gets the public answer or NXDOMAIN. So always find out which resolver the client uses before you debug anything else.

In an interview Junior

How would you move a service to a new IP with DNS, without users hitting the old one for an hour?

Plan around the TTL: every resolver keeps the old answer until its TTL runs out, and you cannot flush someone else's cache.

  1. Check the current TTL on the authoritative server: dig +noall +answer @ns1.lab orders.lab.
  2. Lower it (to 60 or 300) at least one full old TTL before the change - the lowering itself only spreads as the old TTL expires.
  3. Change the record (and bump the SOA serial, or secondaries never pick it up).
  4. Keep the old server running until the old TTL plus some margin has passed: programs that cache names themselves (nginx at startup, Java) hold them longer.
  5. Raise the TTL again afterwards.

Also watch negative caching: if clients looked the new name up before it existed, the NXDOMAIN is cached for the SOA minimum.

Also asked: What is the difference between an A record and a CNAME? · Why can you not put a CNAME at the zone apex? · What is split-horizon DNS, and how can it break a connection?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.