Why this matters
Programs connect to names like api.lab, but packets need IP addresses. Turning one into the other goes through several layers on the box, and each one can answer, remember, or lie. "The app cannot resolve X" is a daily ticket; the fix is always finding which layer answered.
What you need to know already: Ch 1 · Network identity (1.9) - DHCP gave your box an address, a gateway and a DNS server. Ch 4 - symlinks. Ch 2 - systemd services.
A few words first
- DNS (Domain Name System) is the internet's phone book: a worldwide system of servers that turn names into addresses.
- To resolve a name is to look up its address. A resolver is the software (or server) that does the looking up.
- A domain is a name in that system, read right to left:
api.github.comisapiinsidegithubinsidecom.
The path of one lookup
When curl http://api.lab/ needs an address, nothing talks to a DNS server directly at first. The chain on this box (and every Ubuntu server since 18.04):
curl -> getaddrinfo() the C library's lookup function (glibc)
-> /etc/nsswitch.conf "hosts: files dns"
-> files = /etc/hosts checked FIRST
-> dns = /etc/resolv.conf -> nameserver 127.0.0.53
-> systemd-resolved a local cache and forwarder
-> 10.64.0.1 upstream resolver learned from DHCP (your Mac)
-> the recursive chain root -> TLD -> the zone's own servers
getaddrinfo() is the function in glibc (the C library almost every program on Linux is built on) that programs call to turn a name into addresses. The rest of this lesson walks the chain one layer at a time.
nsswitch.conf decides the order
$ grep hosts /etc/nsswitch.conf
hosts: files dns
/etc/nsswitch.conf tells glibc where to look for things. For host names: files first, then dns. So an entry in /etc/hosts wins over DNS, always, for every program that uses the C library - which is nearly all of them (Java, Python, Node, curl, ssh). Some desktop installs have files mdns4_minimal [NOTFOUND=return] dns; the idea is the same.
$ cat /etc/hosts
127.0.0.1 localhost
127.0.1.1 oncall-lab
# The following lines are desirable for IPv6 capable hosts
::1 ip6-localhost ip6-loopback
...
/etc/hosts is a plain file: an address, then names for it. 127.0.1.1 oncall-lab is Ubuntu's way of making the machine's own hostname resolve without DNS.
getent: ask exactly like a program
$ getent hosts api.lab
10.0.3.20 api.lab
$ getent hosts nope.lab; echo "exit $?"
exit 2
getent hosts NAME looks the name up through nsswitch, exactly the way programs do. It is the answer to "what will my service connect to". Output: address, then name. Exit code 2 ($?, Ch 6) = not found. getent ahostsv4 shows every IPv4 address the lookup returns, once per kind of connection (you can ignore the STREAM/DGRAM/RAW column for now):
$ getent ahostsv4 example.com
93.184.216.34 STREAM example.com
93.184.216.34 DGRAM
93.184.216.34 RAW
resolv.conf is a symlink, and 127.0.0.53 is not a real DNS server
$ ls -l /etc/resolv.conf
lrwxrwxrwx 1 root root 39 Sep 3 10:12 /etc/resolv.conf -> ../run/systemd/resolve/stub-resolv.conf
$ cat /etc/resolv.conf
# This is /run/systemd/resolve/stub-resolv.conf managed by man:systemd-resolved(8).
nameserver 127.0.0.53
options edns0 trust-ad
search .
/etc/resolv.conf tells glibc which DNS server to ask (nameserver). On Ubuntu it is a symlink (the l at the start of the mode, and the ->) into a file that systemd-resolved generates. systemd-resolved is a systemd service (like the ones in Ch 2) that does DNS for the whole box. 127.0.0.53 is its stub listener: an address on this box where it waits for questions. glibc sends it the question; resolved answers from its cache (its memory of recent answers) or forwards the question upstream.
Because the file is generated, editing it by hand is useless (it gets replaced), and replacing the symlink with a regular file takes resolved out of the path - the cause of one of this chapter's incidents.
search . means "no search domains" (8.27 explains search domains). options are tuning flags you can ignore for now.
Where resolved actually sends questions
resolvectl is resolved's control command. resolvectl status shows its settings:
$ resolvectl status
Global
Protocols: -LLMNR -mDNS -DNSOverTLS DNSSEC=no/unsupported
resolv.conf mode: stub
Link 2 (enp0s1)
Current Scopes: DNS
Protocols: +DefaultRoute -LLMNR -mDNS -DNSOverTLS DNSSEC=no/unsupported
Current DNS Server: 10.64.0.1
DNS Servers: 10.64.0.1
resolv.conf mode: stub- the symlink is intact.foreignmeans someone replaced it with their own file; resolved is no longer in the path.Link 2 (enp0s1)- settings per interface ("link").DNS Servers- the upstream servers resolved forwards to, learned from DHCP here: your Mac.+DefaultRoute- this link's servers get questions for any domain.- The
Protocolslines list optional features, all off here; ignore them.
$ resolvectl query example.com
example.com: 93.184.216.34 -- link: enp0s1
-- Information acquired via protocol DNS in 23.4ms.
-- Data is authenticated: no; Data was acquired via local or encrypted transport: no
-- Data from: network
$ resolvectl query example.com
...
-- Data from: cache
resolvectl query NAME resolves through resolved and says where the answer came from: network (it asked upstream) or cache (it remembered). Clear the cache with sudo resolvectl flush-caches; see hits and misses with resolvectl statistics.
One subtlety that matters later: resolved reads /etc/hosts too and answers from it, even on 127.0.0.53. It also makes up a few names itself: the local hostname, localhost, and _gateway (your default gateway):
$ dig +short _gateway
10.64.0.1
(dig asks a DNS server a question; +short prints only the answer. It gets its own lesson next.) resolved has a second stub at 127.0.0.54 that passes questions upstream without those local extras.
The recursive chain
Past your resolver, someone has to find the answer starting from nothing. DNS is a tree: the root (.) at the top, then TLDs (top-level domains: com, org, uk...), then each domain. The part of the tree one set of servers is responsible for is a zone.
resolver -> a root server (.) "I don't know api.github.com, but .com is
served by a.gtld-servers.net ..."
-> a .com server "github.com is served by dns1.p08.nsone.net"
-> github.com's own servers "api.github.com is 140.82.121.5, TTL 60"
A resolver that walks this chain for you is a recursive resolver. The first two answers are referrals ("ask them instead"), not answers. Only the last server is authoritative - it holds the real data for that zone. Every resolver along the way caches what it learned for the record's TTL (time to live: how many seconds an answer may be remembered). That is why the second lookup is fast, and why a change takes time to reach everyone.
Caches, plural
program some cache names themselves: Java services (like orders from
Ch 3) keep answers for 30 s by default; nginx looks up the
names in its config ONCE at startup; Go and Node do not cache
resolved per-box cache, respects TTLs (capped at 2 hours)
upstream your router, company resolvers, public ones like 8.8.8.8
The nginx one causes real outages: a server behind a name changes IP, and nginx keeps sending to the old address until it is reloaded, because it looked the name up once when it started.
The tools, and what each one skips
getent hosts X nsswitch: /etc/hosts, then DNS. What your program sees.
dig X DNS only, to the server in resolv.conf (the stub).
Does not read /etc/hosts itself - but on Ubuntu the stub
it asks does.
dig @10.0.3.53 X DNS only, straight to that server. Skips the stub,
its cache and /etc/hosts.
host X / nslookup X DNS only, like dig (older tools, 8.18)
resolvectl query X through resolved, says cache or network
"dig works but the app does not" is the classic symptom of /etc/hosts or nsswitch, and the rule is simple: when a program cannot resolve something, run getent first. When getent and dig @the-real-server disagree, the answer is on this box, not in DNS.
What you can now do
- Name every layer a lookup passes through on Ubuntu, in order.
- Use
getentto see what a program gets, andresolvectlto see resolved's upstream and cache. - Explain referrals, authoritative servers and TTLs.