Why one lookup can cost eight
You met ndots in 8.27 on the Ubuntu box. Inside a pod it bites harder: Kubernetes sets ndots:5, so almost every external name (api.github.com, a payment provider, a database host) is first tried with every cluster search domain. At a few thousand requests per second that is real load on CoreDNS and real latency for your app. This lesson counts the queries and shows the fixes.
What you need to know already: search domains and the ndots rule (8.27), A and AAAA records, NXDOMAIN and NOERROR (8.22), resolv.conf inside a pod, CoreDNS, the Corefile and its reload plugin (16.13), kubectl logs (15.15).
The rule, again
The resolvers in the C libraries programs use (glibc on Ubuntu, musl on Alpine and busybox) decide what to try from ndots: if the name has fewer dots than ndots, try it with every search domain appended first, and only then as written.
Kubernetes sets ndots:5 so that web, web.shop and web.shop.svc all resolve through the search list.
For a pod in shop looking up api.github.com (2 dots, fewer than 5):
1 api.github.com.shop.svc.cluster.local NXDOMAIN
2 api.github.com.svc.cluster.local NXDOMAIN
3 api.github.com.cluster.local NXDOMAIN
4 api.github.com NOERROR 140.82.121.5
Three wasted lookups before the real one. And applications ask for A and AAAA (IPv4 and IPv6) at the same time, so that is 8 queries for one name, 6 of them wasted - per connection, for clients that do not cache. Each miss is a network round trip.
Seeing it
CoreDNS's log plugin prints one line per query it answers. To turn it on, add a line log inside the server block of the Corefile. kubectl edit cm coredns -n kube-system opens the ConfigMap in an editor and saves it back when you quit (15.40). Then wait for the reload plugin to pick it up (up to about 90 s, 16.13):
# after adding `log` to the Corefile (the ndots mission does)
k logs -n kube-system -l k8s-app=kube-dns | grep -E 'Reloading|github'
[INFO] Reloading
[INFO] plugin/reload: Running configuration SHA512 = 3d8f4ebee6bb3333...
[INFO] Reloading complete
[INFO] 10.244.2.16:43589 - 54034 "A IN api.github.com.shop.svc.cluster.local. udp 55 false 512" NXDOMAIN qr,aa,rd 131 0.00012s
[INFO] 10.244.2.16:43104 - 14439 "A IN api.github.com.svc.cluster.local. udp 50 false 512" NXDOMAIN qr,aa,rd 126 0.00016s
[INFO] 10.244.2.16:42716 - 42763 "A IN api.github.com.cluster.local. udp 46 false 512" NXDOMAIN qr,aa,rd 122 0.00008s
[INFO] 10.244.2.16:41358 - 51897 "A IN api.github.com. udp 32 false 512" NOERROR qr,rd,ra 48 0.0284s
The first three lines are the reload. Reading one query line, left to right:
10.244.2.16:43589 the client pod's IP and port
54034 the query id
"A IN api.github... ..." record type, class (IN = internet), name, protocol,
request size, a DNSSEC flag, the client's buffer size
NXDOMAIN the answer code (8.22)
qr,aa,rd flags: aa = authoritative (CoreDNS answered from the
cluster zone itself); ra = recursion available (the
answer came from upstream)
131 response size in bytes
0.00012s how long it took
Look at the last column: the cluster misses take a tenth of a millisecond, the real answer 28 ms (it went upstream). With two CoreDNS replicas the lines are split between their logs; -l k8s-app=kube-dns reads both.
The log plugin logs every query. On a busy cluster that is a lot of log volume and some CPU: turn it on to investigate, off afterwards.
Fix 1: a trailing dot
A trailing dot makes the name absolute (8.27): no search list at all.
$ k exec -n shop toolbox -- nslookup api.github.com.
Name: api.github.com
Address: 140.82.121.5
2 queries instead of 8. It works everywhere a hostname is accepted - but some HTTP clients then send Host: api.github.com. (with the dot, 9.21) and some TLS libraries refuse the certificate for the dotted name (9.15). Test it per client.
Fix 2: lower ndots for the pod
dnsConfig (16.13) overrides the option for one pod:
spec:
dnsConfig:
options:
- name: ndots
value: "2"
# payments = a Deployment with the dnsConfig above (the ndots mission)
k exec deploy/payments -- cat /etc/resolv.conf
search payments.svc.cluster.local svc.cluster.local cluster.local
nameserver 10.96.0.10
options ndots:2
Now api.github.com (2 dots, not fewer than 2) is tried as written first.
The cost: web.shop.svc (also 2 dots) is also tried as written first - one wasted query to the upstream server, which returns NXDOMAIN for the made-up top-level domain svc, before the search list saves it. Short names like web and web.shop still work. ndots:1 or 2 is the usual compromise for pods that mostly call external APIs.
Other levers
You will meet these in real clusters:
- NodeLocal DNSCache: a small caching DNS server on every node (a DaemonSet). Pods ask their own node, so there is no network hop. It also avoids a known conntrack (16.3) race on busy nodes that makes some UDP lookups vanish and wait out a 5-second timeout.
- CoreDNS autopath: the server works out the right name on the first query and answers it directly; it needs
pods verifiedin the kubernetes plugin, which costs memory. - Caching in the application. Some language runtimes cache DNS answers themselves, sometimes for a surprisingly long or short time - check yours.
Later (Ch 20): how the JVM caches DNS answers, and the setting that controls it.
Numbers to have in your head
name with N dots, ndots:5, pod in namespace X queries (A+AAAA)
web (0 dots) -> found at step 1 2
web.shop (1 dot) -> found at step 2 4
api.github.com (2 dots) -> found at step 4 8
api.github.com. (absolute) 2
What you can now do:
- count the DNS queries behind one lookup from a pod
- turn on CoreDNS query logging and read a log line
- cut the waste with a trailing dot or a per-pod ndots, knowing each one's cost