Why this matters
A service moves to a new server, DNS is changed, and half the users keep hitting the old one for an hour. Or you create a new name and it "does not exist" for five minutes. Both follow directly from how records and TTLs work - and both are avoidable if you plan.
What you need to know already: 8.16 (zones, authoritative servers, TTL) and 8.18 (reading dig).
The record types
A zone holds records: one fact each (a name, a type, a TTL, the data):
A name -> IPv4 address
AAAA name -> IPv6 address
CNAME name -> another name (an alias). The resolver restarts the lookup.
NS which servers are authoritative for a zone (and for delegations)
SOA one per zone: primary server, admin, serial, timers, negative TTL
MX mail servers with a priority (LOWER number = preferred)
TXT free text: used to prove you own a domain, and for mail settings
SRV where a service runs: priority weight port target
PTR IP -> name, under in-addr.arpa (reverse DNS)
CAA which companies may issue security certificates for the name
An SRV record, read field by field:
_postgresql._tcp.db.lab. 300 IN SRV 0 5 5432 db.lab.
| | | target host
| | port
| weight (among same priority)
priority (lower first)
The name says which service (_postgresql, a database) over which transport (_tcp), in which domain.
CNAME has rules, and they bite
- A CNAME cannot sit next to any other record at the same name. If
wwwis a CNAME, it cannot also have TXT or MX records. - No CNAME at the zone apex. The apex is the zone's own name,
example.comitself. It must have SOA and NS records, so by rule 1 it cannot be a CNAME. You cannot pointexample.comat a hosting company's name with a CNAME.
That is why DNS providers invented their own workarounds ("alias" or "flattened" records): their server looks up the target itself and hands out plain A records.
A CNAME chain costs a lookup per hop and every hop has its own TTL. The answer's lifetime is the shortest TTL in the chain.
The SOA record
lab. 3600 IN SOA ns1.lab. hostmaster.lab. 2026092301 7200 3600 1209600 300
| | | | | | |
primary admin (hostmaster@lab) | | | minimum
serial | | expire
refresh retry
- primary - the main server for the zone. admin - a contact address, with the first dot standing for @.
- serial - a version number. Bump it on every change, or the backup (secondary) servers never pick up the change.
YYYYMMDDnnis the convention. - refresh / retry / expire - how often secondaries check the primary for a new serial.
- minimum - today it means one thing: how long a "no such name" answer is cached.
TTL: why your change did not take effect
Every resolver that holds the old answer keeps it until its TTL runs out. You cannot flush someone else's cache. So a change to a record with TTL 3600 takes up to an hour to reach everyone, and some programs (Java services, nginx, 8.16) hold it longer.
The migration procedure:
T-48h check the current TTL: dig +noall +answer @ns1.lab orders.lab
T-24h lower it to 60 (or 300)
lowering a TTL is itself subject to the OLD TTL - resolvers that cached
the record keep the 3600 until it expires. So do it a full old-TTL
(plus margin) before the change, not an hour before.
T-0 change the record
T+5m verify from several resolvers; watch the old server's logs
T+1d raise the TTL back (60s forever = 60x the query load, for nothing)
(T-24h = 24 hours before the change, T+5m = 5 minutes after.)
Negative caching: NXDOMAIN sticks too
When a name does not exist, resolvers remember that as well - negative caching - for:
negative TTL = min(SOA record's own TTL, SOA minimum field)
For lab. that is min(3600, 300) = 300 seconds. The classic trap:
- The app starts and looks up
newsvc.lab- it does not exist yet: NXDOMAIN. - You create the record a minute later.
- For up to five more minutes, that app's resolver keeps answering NXDOMAIN.
"I created the record and it still says it does not exist" is almost always this. On your own box, sudo resolvectl flush-caches fixes it; on everyone else's, you wait.
Split-horizon DNS
The same name gives different answers depending on who asks - split horizon. Company resolvers serve an internal view; the internet sees a public one, or nothing:
$ dig +short api.lab # from inside: the internal resolver
10.0.3.20
$ dig +short @8.8.8.8 api.lab # from the internet's point of view
$ dig @8.8.8.8 api.lab | grep status
;; ->>HEADER<<- opcode: QUERY, status: NXDOMAIN, id: 9912
(8.8.8.8 is Google's public resolver - it only knows the public internet.) It is deliberate, and common:
- Company DNS:
intranet.company.comresolves only inside the office. - Private addresses for public names: a company can make a service's public name resolve to a private 10.x address for machines inside its network, and to the public address for everyone else.
The classic failure: a machine is set up to use a public resolver (or any resolver that does not have the internal view). It gets the public address, connects there, and is refused, because the service only accepts connections on the private one. The network looks fine; DNS gave the wrong view.
When you debug DNS across networks, always ask: which resolver did this client use, and which view does that resolver have?
Later (Ch 23): cloud "private endpoints" work exactly like this, and the wrong-resolver failure is one of the most common cloud tickets.
What you can now do
- Name the common record types and read an SRV and an SOA record.
- Plan a DNS change around its TTL.
- Explain negative caching and split horizon, and the questions each one raises.