OnCallReady

Lesson 8.27 · Addressing & DNS · 12 min read

Search domains and the ndots trap

In plain words

Imagine a school where, if you ask the receptionist for "Sam", she first checks "Sam in your class", then "Sam in your year", then "Sam in the whole school", and only then looks outside the school. That is handy: you can say "Sam" instead of "Sam Jones, class 4B".

But the receptionist has a rule: any name with fewer than five dots gets the full search first. So when you ask for api.github.com, she dutifully checks three made-up school names before trying the real one. That is a resolv.conf with search domains and options ndots:5: short names work, and every outside name pays for it with extra queries, twice over because glibc asks for both A and AAAA.

Why this matters

Inside a company you type ssh db instead of ssh db.prod.company.com - a resolv.conf setting makes the short name work. The same setting can make every lookup of an outside name cost four times the DNS traffic. This lesson shows the rule, lets you count the queries, and lists the fixes.

What you need to know already: 8.16 (resolv.conf, the stub), 8.18 (dig, NXDOMAIN, A and AAAA), 8.14 (tcpdump). Environment variables from Ch 6.

Search domains

resolv.conf can list search domains: suffixes the resolver tries appending to names that are not written in full. A config you will meet on many platforms:

search default.svc.cluster.local svc.cluster.local cluster.local
nameserver 10.96.0.10
options ndots:5

The rule

ndots ("number of dots"): if the name has fewer than ndots dots, try it with every search suffix first, and only then as written. If it has ndots or more, try it as written first, then the search list. A name ending in . is absolute (fully qualified, 8.18): no search at all.

api.github.com has 2 dots. 2 < 5, so the resolver tries:

1. api.github.com.default.svc.cluster.local.   NXDOMAIN
2. api.github.com.svc.cluster.local.           NXDOMAIN
3. api.github.com.cluster.local.               NXDOMAIN
4. api.github.com.                             the answer

And glibc asks for A and AAAA for each one (both at once), so that is eight queries to resolve one outside name, six of them certain to fail. Add more search domains and it gets worse.

At 200 requests a second to an outside service, with a fresh lookup for each request, that is 1,600 DNS queries a second hitting the internal DNS server. It shows up as that server's CPU climbing, and as slow calls to the outside service.

Why would anyone set it up this way? So that short names work: with that search list, orders becomes orders.default.svc.cluster.local on the first try. Short names are cheap; outside names pay.

Later (Ch 16): this exact resolv.conf is what every app gets on a Kubernetes cluster, which is why this trap is famous.

See it on this box

glibc lets you override resolv.conf for one command with two environment variables - perfect for experiments:

LOCALDOMAIN   replaces the search list
RES_OPTIONS   adds options, e.g. ndots:5

Putting VAR=value in front of a command sets it for that command only (Ch 6). Capture port 53 in the background (-i lo = the loopback interface, where programs talk to the stub; -c 16 = stop after 16 packets; port 53 = only DNS), then resolve with the search list:

$ sudo tcpdump -i lo -nn -c 16 port 53 &
$ LOCALDOMAIN='default.svc.cluster.local svc.cluster.local cluster.local' RES_OPTIONS=ndots:5 getent hosts api.github.com
10:12:01.001200 IP 127.0.0.1.40112 > 127.0.0.53.53: 3321+ A? api.github.com.default.svc.cluster.local. (58)
10:12:01.002900 IP 127.0.0.53.53 > 127.0.0.1.40112: 3321 NXDomain 0/1/0 (108)
10:12:01.003100 IP 127.0.0.1.40113 > 127.0.0.53.53: 8812+ AAAA? api.github.com.default.svc.cluster.local. (58)
...
10:12:01.042100 IP 127.0.0.53.53 > 127.0.0.1.40117: 1440 1/0/0 A 140.82.121.5 (48)
140.82.121.5    api.github.com

Reading a tcpdump DNS line: time, IP, sender.port > receiver.port, the query id (+ = recursion wanted), the question (A? = "what is the A record of...") or the reply (NXDomain, or 1/0/0 = one answer record, then the answer). The last line is getent's own output.

Sixteen packets: eight questions, eight answers. On -i lo you see the program talking to the stub; with -i any (every interface) you would also see resolved's own questions to the upstream on enp0s1 (only for names it has not cached). dig can show the search walk too, but only if you ask - by default dig ignores the search list:

$ LOCALDOMAIN='default.svc.cluster.local svc.cluster.local' RES_OPTIONS=ndots:5 dig +search +showsearch +noall +comments api.github.com | grep status
;; ->>HEADER<<- opcode: QUERY, status: NXDOMAIN, id: 101
;; ->>HEADER<<- opcode: QUERY, status: NXDOMAIN, id: 102
;; ->>HEADER<<- opcode: QUERY, status: NOERROR, id: 103

(+search = use the search list; +showsearch = print every try; +comments = keep the header lines, so grep can find status.)

The fixes, cheapest first

1. The trailing dot. A fully qualified name skips the search list:

curl https://api.github.com./repos/...

Two queries instead of eight. Free. The catch: web clients and servers compare the name you used against what the server expects, and some reject the trailing dot. Test it before you rely on it.

2. Lower ndots for that program. With ndots:2, api.github.com (2 dots) is tried as written first. Short names like orders (0 dots) still use the search list.

3. A cache close by - a small caching resolver on each machine, so the doomed queries are answered locally instead of crossing the network. It fixes the load, not the waste.

4. Reuse connections. A program that keeps its connection open resolves once per connection, not once per request. Often the real fix.

Before you add more DNS servers

When a DNS server is overloaded, look first at what it answers: if most answers are NXDOMAIN for outside names with your search suffixes glued on, this is ndots, and more servers just pay for the waste in more places.

What you can now do

Why it helps

This is a classic performance problem, and a favourite interview topic, because the symptoms are indirect: the internal DNS server's CPU climbing, slow calls to outside services, lots of NXDOMAIN answers nobody asked for. Teams often respond by adding DNS servers, which pays for the wasted queries in more places.

Knowing the mechanism lets you prove it (count the queries with tcpdump, look at the share of NXDOMAIN answers for outside names with your search suffixes glued on) and pick the right fix: a trailing dot, a lower ndots for the chatty program, a local cache, or simply reusing connections. It also explains why host db works while dig db does not.

Commands in this lesson

tcpdump

FAQ

Why would anyone set ndots:5?

So that short names work. With search default.svc.cluster.local svc.cluster.local cluster.local, a program can use orders and it becomes orders.default.svc.cluster.local on the first try; even names with a few dots, like orders.shop.svc, go through the search list. The cost is that every outside name with fewer than five dots is tried with each suffix first.

How many queries does one outside lookup really cost?

With three search domains and ndots:5, api.github.com is tried three times with a suffix (all NXDOMAIN) and once as written. glibc asks for A and AAAA each time, so eight queries, six of them certain to fail. More search domains make it worse. Multiply by requests per second when the program looks the name up again for every request.

Is the trailing dot safe to use everywhere?

The DNS side is: api.github.com. is absolute and skips the search list, two queries instead of eight. The web side is not always. The program may send the name with the dot to the server, and some servers or security checks compare names exactly and reject it. Most clients strip the dot, some do not. Test it with the real program and service before relying on it.

Does lowering ndots break short internal names?

Mostly not. With ndots:2, names with 0 or 1 dots (orders, orders.shop) still go through the search list first. Names with 2 or more dots are tried as written first: outside names get faster, and internal names like orders.shop.svc fail once before the search list rescues them. Full names are unaffected. Change it only for the program that needs it.

Why use LOCALDOMAIN and RES_OPTIONS instead of editing resolv.conf?

Because they change the behaviour for one command only, and resolv.conf on Ubuntu belongs to systemd-resolved (lesson 8.16). LOCALDOMAIN replaces the search list and RES_OPTIONS adds options for the command you prefix them to, so you can reproduce any client's setup on your own box, count queries, and leave the box exactly as it was.

In an interview Junior

A resolv.conf has "search default.svc.cluster.local svc.cluster.local cluster.local" and "options ndots:5". What happens when a program looks up api.github.com?

The ndots rule: a name with fewer than ndots dots is tried with every search domain appended first, and only then as written.

api.github.com has 2 dots, 2 < 5, so the resolver tries:

  1. api.github.com.default.svc.cluster.local - NXDOMAIN
  2. api.github.com.svc.cluster.local - NXDOMAIN
  3. api.github.com.cluster.local - NXDOMAIN
  4. api.github.com - the answer

glibc asks for A and AAAA for each, so that is eight queries for one outside name, six certain to fail. At high request rates that is load on the internal DNS server and latency on every call.

Fixes, cheapest first: a trailing dot (api.github.com., fully qualified, no search), a lower ndots for that program, a cache close by, and reusing connections. You can count the queries with tcpdump -i lo port 53.

Also asked: What is a search domain in resolv.conf? · What does a trailing dot at the end of a hostname mean? · Why does host find a short name that dig says is NXDOMAIN?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.