OnCallReady

Lesson 16.14 · Kubernetes: Networking & Storage · 15 min read

ndots:5, the search path, and why api.github.com costs 8 queries

In plain words

Imagine a rule at school: "if a name has fewer than five dots, first assume it's a nickname and look for it in your class, your year and your school, and only then treat it as a full address". Ask for "Maria" and it's found in your class right away. But ask for "api.github.com" and you first search three class lists that obviously won't have it, and you do each search twice (for two kinds of address).

That rule is options ndots:5 in every pod's /etc/resolv.conf. api.github.com has two dots, so the resolver tries api.github.com.shop.svc.cluster.local, .svc.cluster.local and .cluster.local first, all NXDOMAIN, for both A and AAAA. Eight queries for one name. A trailing dot or a lower ndots fixes it.

Why one lookup can cost eight

You met ndots in 8.27 on the Ubuntu box. Inside a pod it bites harder: Kubernetes sets ndots:5, so almost every external name (api.github.com, a payment provider, a database host) is first tried with every cluster search domain. At a few thousand requests per second that is real load on CoreDNS and real latency for your app. This lesson counts the queries and shows the fixes.

What you need to know already: search domains and the ndots rule (8.27), A and AAAA records, NXDOMAIN and NOERROR (8.22), resolv.conf inside a pod, CoreDNS, the Corefile and its reload plugin (16.13), kubectl logs (15.15).

The rule, again

The resolvers in the C libraries programs use (glibc on Ubuntu, musl on Alpine and busybox) decide what to try from ndots: if the name has fewer dots than ndots, try it with every search domain appended first, and only then as written.

Kubernetes sets ndots:5 so that web, web.shop and web.shop.svc all resolve through the search list.

For a pod in shop looking up api.github.com (2 dots, fewer than 5):

1  api.github.com.shop.svc.cluster.local    NXDOMAIN
2  api.github.com.svc.cluster.local         NXDOMAIN
3  api.github.com.cluster.local             NXDOMAIN
4  api.github.com                           NOERROR   140.82.121.5

Three wasted lookups before the real one. And applications ask for A and AAAA (IPv4 and IPv6) at the same time, so that is 8 queries for one name, 6 of them wasted - per connection, for clients that do not cache. Each miss is a network round trip.

Seeing it

CoreDNS's log plugin prints one line per query it answers. To turn it on, add a line log inside the server block of the Corefile. kubectl edit cm coredns -n kube-system opens the ConfigMap in an editor and saves it back when you quit (15.40). Then wait for the reload plugin to pick it up (up to about 90 s, 16.13):

# after adding `log` to the Corefile (the ndots mission does)
k logs -n kube-system -l k8s-app=kube-dns | grep -E 'Reloading|github'
[INFO] Reloading
[INFO] plugin/reload: Running configuration SHA512 = 3d8f4ebee6bb3333...
[INFO] Reloading complete
[INFO] 10.244.2.16:43589 - 54034 "A IN api.github.com.shop.svc.cluster.local. udp 55 false 512" NXDOMAIN qr,aa,rd 131 0.00012s
[INFO] 10.244.2.16:43104 - 14439 "A IN api.github.com.svc.cluster.local. udp 50 false 512" NXDOMAIN qr,aa,rd 126 0.00016s
[INFO] 10.244.2.16:42716 - 42763 "A IN api.github.com.cluster.local. udp 46 false 512" NXDOMAIN qr,aa,rd 122 0.00008s
[INFO] 10.244.2.16:41358 - 51897 "A IN api.github.com. udp 32 false 512" NOERROR qr,rd,ra 48 0.0284s

The first three lines are the reload. Reading one query line, left to right:

10.244.2.16:43589          the client pod's IP and port
54034                      the query id
"A IN api.github...  ..."  record type, class (IN = internet), name, protocol,
                           request size, a DNSSEC flag, the client's buffer size
NXDOMAIN                   the answer code (8.22)
qr,aa,rd                   flags: aa = authoritative (CoreDNS answered from the
                           cluster zone itself); ra = recursion available (the
                           answer came from upstream)
131                        response size in bytes
0.00012s                   how long it took

Look at the last column: the cluster misses take a tenth of a millisecond, the real answer 28 ms (it went upstream). With two CoreDNS replicas the lines are split between their logs; -l k8s-app=kube-dns reads both.

The log plugin logs every query. On a busy cluster that is a lot of log volume and some CPU: turn it on to investigate, off afterwards.

Fix 1: a trailing dot

A trailing dot makes the name absolute (8.27): no search list at all.

$ k exec -n shop toolbox -- nslookup api.github.com.
Name:	api.github.com
Address: 140.82.121.5

2 queries instead of 8. It works everywhere a hostname is accepted - but some HTTP clients then send Host: api.github.com. (with the dot, 9.21) and some TLS libraries refuse the certificate for the dotted name (9.15). Test it per client.

Fix 2: lower ndots for the pod

dnsConfig (16.13) overrides the option for one pod:

spec:
  dnsConfig:
    options:
    - name: ndots
      value: "2"
# payments = a Deployment with the dnsConfig above (the ndots mission)
k exec deploy/payments -- cat /etc/resolv.conf
search payments.svc.cluster.local svc.cluster.local cluster.local
nameserver 10.96.0.10
options ndots:2

Now api.github.com (2 dots, not fewer than 2) is tried as written first.

The cost: web.shop.svc (also 2 dots) is also tried as written first - one wasted query to the upstream server, which returns NXDOMAIN for the made-up top-level domain svc, before the search list saves it. Short names like web and web.shop still work. ndots:1 or 2 is the usual compromise for pods that mostly call external APIs.

Other levers

You will meet these in real clusters:

Later (Ch 20): how the JVM caches DNS answers, and the setting that controls it.

Numbers to have in your head

name with N dots, ndots:5, pod in namespace X        queries (A+AAAA)
web               (0 dots)   -> found at step 1       2
web.shop          (1 dot)    -> found at step 2       4
api.github.com    (2 dots)   -> found at step 4       8
api.github.com.   (absolute)                          2

What you can now do:

Why it helps

This shows up as a real incident: a service calling an external payment or fraud API at a few thousand requests per second triples CoreDNS load, adds latency to every call, and occasionally hits 5-second DNS timeouts on busy nodes. The app team sees "slow partner API"; the actual cause is six wasted DNS queries per connection.

Knowing it lets you read CoreDNS log plugin output and spot the NXDOMAIN pattern immediately, and pick the right fix for the workload: dnsConfig with ndots: "2", a trailing dot, NodeLocal DNSCache or app-side caching. It's also a favourite senior interview question, because it tests whether you understand resolvers, not just Kubernetes.

FAQ

Why does Kubernetes use ndots:5 if it wastes queries?

So that every short in-cluster form resolves through the search list: web, web.shop, web.shop.svc, and even names like web-0.web.shop.svc with more dots. A high ndots guarantees cluster names are tried as cluster names first. The price is paid by external names, which go through every search domain before being tried as written.

Does a trailing dot always work?

DNS-wise, yes: api.github.com. is absolute, the search list is skipped, and you get 2 queries instead of 8. But some HTTP clients then send Host: api.github.com. and some TLS stacks refuse the certificate for the dotted name, and URLs with trailing dots look odd in config. Test it per client before rolling it out.

What breaks if I set ndots to 1 or 2?

Short names still work: web (0 dots) and web.shop (1 dot) go through the search list. Names with at least as many dots as ndots are tried as written first, so with ndots:2 web.shop.svc costs one extra upstream query (NXDOMAIN for the fake TLD svc) before the search list finds it. Always use web.shop or the full name and it's fine.

Why 8 queries and not 4?

Because modern resolvers ask for the IPv4 (A) and IPv6 (AAAA) records of each candidate name in parallel. Four candidates times two record types is eight. On an IPv4-only cluster the AAAA answers are useless too, but the client still asks unless configured otherwise.

Don't applications cache DNS anyway?

Some do: many language runtimes keep answers for a while (often around 30 seconds, sometimes forever by default). So a long-running process with such a cache repeats the expensive lookup far less often. Short-lived processes, scripts, and clients that open a new connection per request without a resolver cache pay the full cost every time.

In an interview Mid

What does ndots:5 do in a pod's resolv.conf, and why can it be a problem?

The ndots rule: a name with fewer dots than ndots is tried with every search domain appended first, and only then as written. Kubernetes sets ndots:5 so that web, web.shop and web.shop.svc resolve through the search list.

The cost lands on external names. For api.github.com (2 dots) from a pod in shop:

  1. api.github.com.shop.svc.cluster.local - NXDOMAIN
  2. api.github.com.svc.cluster.local - NXDOMAIN
  3. api.github.com.cluster.local - NXDOMAIN
  4. api.github.com - the answer

With A and AAAA asked together, that is 8 queries for one name, 6 wasted - load on CoreDNS and latency per connection. You can see it with CoreDNS's log plugin.

Fixes: a trailing dot (api.github.com., no search; test the client), a lower ndots for the pod through dnsConfig.options, NodeLocal DNSCache, and caching or connection reuse in the app.

Also asked: How would you check how many DNS queries a pod makes for an external name? · What is NodeLocal DNSCache, and what does it fix? · What does a trailing dot on a hostname do?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.