OnCallReady

Lesson 16.13 · Kubernetes: Networking & Storage · 23 min read

CoreDNS and the names it serves

In plain words

Every classroom has a phone book pinned by the door, and the first page says: "if you just write a first name, assume they're in this class; if not, try this school; then this town". So you can say "Maria" and it finds Maria from your class, or say "Maria.room5" to reach another class.

In Kubernetes the kubelet writes that page into each pod's /etc/resolv.conf: the nameserver is the kube-dns Service (10.96.0.10), and the search list tries shop.svc.cluster.local, then svc.cluster.local, then cluster.local. CoreDNS is the operator behind that number: it answers cluster names from the API and forwards everything else to the node's upstream resolver, as configured in its Corefile.

Why the cluster needs its own DNS

In 16.1 a pod reached web, web.shop and web.shop.svc.cluster.local - names nobody ever typed into a DNS server. Services come and go all day; no human could keep a zone file up to date. So the cluster runs its own DNS server that reads Services straight from the API. When "nothing can reach anything", it is often this server, so you need to know how it works.

What you need to know already: how a name becomes an address, resolv.conf, search domains (8.16), dig (8.18), record types A, AAAA, CNAME, PTR, TTLs and NXDOMAIN (8.22), ndots (8.27), Services and ClusterIP (16.1), headless Services and SRV records (16.6, 16.9), ConfigMaps (15.29), CoreDNS named as a cluster add-on (15.7).

Who answers

Every pod's /etc/resolv.conf (8.16) points at one address:

$ k exec -n shop toolbox -- cat /etc/resolv.conf
search shop.svc.cluster.local svc.cluster.local cluster.local
nameserver 10.96.0.10
options ndots:5

Line by line: search = the search domains, tried in order for short names; nameserver = the DNS server to ask; options ndots:5 = the rule from 8.27 (next lesson).

The kubelet writes that file when it starts the pod. Its config holds clusterDNS: [10.96.0.10] (the server) and clusterDomain: cluster.local (the cluster domain: the DNS suffix all cluster names end in).

10.96.0.10 is the ClusterIP of Service kube-dns in kube-system. The name is historical (the older DNS server was called kube-dns); behind it run the CoreDNS pods. get svc,deploy lists both kinds in one command, and -l k8s-app=kube-dns picks them by label:

$ k get svc,deploy -n kube-system -l k8s-app=kube-dns
NAME               TYPE        CLUSTER-IP   EXTERNAL-IP   PORT(S)                  AGE
service/kube-dns   ClusterIP   10.96.0.10   <none>        53/UDP,53/TCP,9153/TCP   13d
NAME                      READY   UP-TO-DATE   AVAILABLE   AGE
deployment.apps/coredns   2/2     2            2           13d

Port 53 over UDP and TCP is DNS; 9153 is where CoreDNS publishes its own counters (queries served, errors) for monitoring tools. The Deployment runs 2 CoreDNS pods, both ready.

So DNS is just another Service: queries go through kube-proxy's rules (16.3) to one of the CoreDNS pods, NetworkPolicies (16.29) apply to them, and if CoreDNS is down, every name lookup in the cluster fails at once.

The records

CoreDNS answers these names. For Service web in namespace shop (cluster domain cluster.local):

web.shop.svc.cluster.local                     A    the ClusterIP
_http._tcp.web.shop.svc.cluster.local          SRV  0 100 80 web.shop.svc.cluster.local.
                                                    (only for NAMED ports)
headless Service:  web.shop.svc.cluster.local  A    one record per ready pod IP
StatefulSet pod:   web-0.web.shop.svc.cluster.local   A  that pod (hostname.subdomain.ns.svc...)
any pod:           10-244-1-16.shop.pod.cluster.local A  10.244.1.16 (the IP with dashes)
ExternalName:      db.shop.svc.cluster.local   CNAME the external name
reverse:           147.142.96.10.in-addr.arpa  PTR  web.shop.svc.cluster.local.

The pattern is always <service>.<namespace>.svc.<cluster domain>. The svc part says "this is a Service" (pods live under pod). The last line is a reverse lookup (8.22): IP to name.

nslookup NAME asks the configured server and prints the answer. The first two lines say which server answered; Name / Address is the result. Given an IP, nslookup does the reverse lookup:

$ k exec -n shop toolbox -- nslookup web
Server:		10.96.0.10
Address:	10.96.0.10:53

Name:	web.shop.svc.cluster.local
Address: 10.96.142.147

$ k exec -n shop toolbox -- nslookup 10.96.142.147
Server:		10.96.0.10
Address:	10.96.0.10:53

147.142.96.10.in-addr.arpa	name = web.shop.svc.cluster.local.

The pod records (10-244-1-16.shop.pod) come from the line pods insecure in CoreDNS's config (below): CoreDNS answers for any dashed IP without checking that such a pod exists. They exist so that a wildcard TLS certificate (*.shop.pod.cluster.local, 9.15) can work; almost nothing else should use them.

Names resolve across namespaces - DNS is not a security boundary:

web                              only from inside namespace shop (search list)
web.shop                         from anywhere ("web" in "shop")
web.shop.svc                     from anywhere
web.shop.svc.cluster.local       from anywhere - the only form that is unambiguous

nslookup, dig and the search list

busybox's nslookup walks the search list and prints every miss. dig does not use the search list unless told to. nicolaka/netshoot is an image full of network tools (dig, curl, tcpdump...), handy as a test pod:

$ k run t -n shop --image=nicolaka/netshoot -- sleep 1d
pod/t created
$ k exec -n shop t -- dig +short web.shop.svc.cluster.local
10.96.142.147
$ k exec -n shop t -- dig web
;; ->>HEADER<<- opcode: QUERY, status: NXDOMAIN, id: 25838
;; QUESTION SECTION:
;web.			IN	A
$ k exec -n shop t -- dig +search +short web.shop
10.96.142.147

+short prints only the answer; +search makes dig use the search list. dig web asked for a top-level domain called web (like com): NXDOMAIN. That is dig being literal, not DNS being broken - a classic false alarm in incident channels.

Test with the full name, or +search, or nslookup / getent hosts NAME. getent hosts asks exactly like applications do - through the C library's resolver (libc), with the search list.

The Corefile

CoreDNS's configuration file is the Corefile. On a cluster it lives in the ConfigMap coredns in kube-system; the jsonpath prints just that key:

$ k get cm coredns -n kube-system -o jsonpath='{.data.Corefile}'
.:53 {
    errors                       # log errors to stdout
    health {                     # :8080/health - the liveness probe
       lameduck 5s
    }
    ready                        # :8181/ready - the readiness probe
    kubernetes cluster.local in-addr.arpa ip6.arpa {
       pods insecure             # answer the 10-244-1-16.ns.pod records
       fallthrough in-addr.arpa ip6.arpa
       ttl 30
    }
    prometheus :9153             # metrics
    forward . /etc/resolv.conf { # everything else: the node's upstream
       max_concurrent 1000
    }
    cache 30 {
       disable success cluster.local
       disable denial cluster.local
    }
    loop                         # detect forwarding loops and exit
    reload                       # re-read this file when it changes
    loadbalance                  # shuffle A records
}

The file is one server block: .:53 { ... } = "for every name (. is the DNS root, so everything), on port 53". Inside, one plugin per line - each plugin is one feature. Order in the file does not matter: CoreDNS runs plugins in a fixed order built into it.

What each one does:

errors        print errors to the pod's log
health        an HTTP page on :8080/health the kubelet checks to see CoreDNS is
              alive; lameduck 5s = keep answering 5 s after being told to stop
ready         an HTTP page on :8181/ready: "I have loaded everything"
kubernetes    answer the cluster zone from the API (below)
(metrics)    the :9153 line publishes CoreDNS's counters for monitoring tools
forward       send every other name to an upstream DNS server
cache 30      remember answers for up to 30 s (but not for cluster.local here)
loop          detect a forwarding loop and stop
reload        re-read the Corefile when it changes
loadbalance   shuffle the order of A records in each answer

The important ones in detail:

dnsPolicy and dnsConfig

A pod's dnsPolicy decides what goes into its resolv.conf:

ClusterFirst            (default) cluster DNS, search list, ndots:5
Default                 the node's resolver - NOT the default, despite the name
ClusterFirstWithHostNet what hostNetwork pods need to use cluster DNS
None                    nothing; you supply everything in dnsConfig

(A hostNetwork pod uses the node's network directly instead of its own, like Docker's --network host, 11.15.)

dnsConfig adds to or overrides the generated file; hostAliases adds lines to the pod's /etc/hosts:

spec:
  dnsPolicy: ClusterFirst
  dnsConfig:
    options:
    - name: ndots
      value: "2"
    searches:
    - corp.lab
  hostAliases:                 # extra /etc/hosts lines, per pod
  - ip: 10.0.3.12
    hostnames: [legacy-db]

dnsConfig is merged into what the policy generates: nameservers appended (max 3), searches appended, options overridden by name. With dnsPolicy: None and no dnsConfig.nameservers, the pod has no nameserver at all.

Debugging DNS from a pod

k run dnsdebug --rm -it --image=busybox:1.36 --restart=Never -- nslookup kubernetes.default
k exec -it deploy/api -- cat /etc/resolv.conf
k get pods -n kube-system -l k8s-app=kube-dns          # running? ready?
k get endpointslice -n kube-system -l kubernetes.io/service-name=kube-dns
k logs -n kube-system -l k8s-app=kube-dns --tail=20

Line by line:

run ... --rm -it --restart=Never   a one-off pod: interactive, never restarted,
                                   deleted when the command ends
exec deploy/api                    exec into any one pod of Deployment api
get pods -l k8s-app=kube-dns       are the CoreDNS pods running and Ready?
get endpointslice ...              does the kube-dns Service have endpoints?
logs ... --tail=20                 the last 20 lines from every CoreDNS pod

nslookup kubernetes.default is the standard test (kubernetes is the Service for the API server, in namespace default): it exercises the search list, CoreDNS, the kubernetes plugin and the API.

What you can now do:

Why it helps

When CoreDNS is down or a NetworkPolicy blocks UDP 53, every service in the cluster fails at once, with errors that look like app bugs: UnknownHostException, Could not resolve host, bad address. Knowing that DNS is just another Service behind kube-proxy lets you check it in three commands: are the CoreDNS pods Ready, does the kube-dns Service have endpoints, does nslookup kubernetes.default work from a pod.

It saves you from the classic false alarm in an incident channel, where someone runs dig web, gets NXDOMAIN and declares DNS broken (dig doesn't use the search list). And on fresh Ubuntu clusters you'll recognise CoreDNS in CrashLoopBackOff with "Loop detected" as the 127.0.0.53 stub resolver problem instead of losing an afternoon.

FAQ

Why is the Service called kube-dns if it runs CoreDNS?

History. Kubernetes shipped a DNS add-on called kube-dns before CoreDNS replaced it (default since 1.13). The Service name and the k8s-app=kube-dns label were kept so that nothing depending on them broke. Behind the kube-dns Service in kube-system run the coredns Deployment's pods.

dig says NXDOMAIN but my app resolves the name fine. Who's right?

Both. dig does not use the search list unless you pass +search, so dig web asks for a top-level domain called web and gets NXDOMAIN. Applications use the libc resolver, which applies the search list. Test with the full name (web.shop.svc.cluster.local), dig +search, nslookup or getent hosts, which behaves like the app.

Why does CoreDNS crash with "Loop detected" on Ubuntu nodes?

CoreDNS forwards non-cluster names to /etc/resolv.conf, and with dnsPolicy: Default that is the node's file. On Ubuntu, /etc/resolv.conf points at the systemd-resolved stub 127.0.0.53, which inside the pod is the pod itself, so CoreDNS would forward to itself. The loop plugin detects that and exits. The fix is the kubelet's resolvConf: /run/systemd/resolve/resolv.conf (the real upstreams), which kubeadm normally sets.

I edited the CoreDNS ConfigMap. When does it take effect?

The reload plugin re-reads the Corefile about every 30 seconds plus jitter, and the kubelet takes up to a minute to sync a ConfigMap into the pod, so plan on up to about a minute and a half. A broken edit is logged and ignored, and the old config keeps running until the pods restart, at which point they crash. Check the logs for "Reloading complete" before walking away.

Is dnsPolicy: Default the default?

No, despite the name. The default is ClusterFirst: cluster DNS, the search list and ndots:5. Default means "use the node's resolver", which is what CoreDNS itself uses. ClusterFirstWithHostNet is what hostNetwork pods need to use cluster DNS, and None gives the pod nothing except what you put in dnsConfig.

In an interview Junior

How does service discovery work in Kubernetes?

Through DNS served by CoreDNS:

So web works from the same namespace, web.shop from anywhere. DNS is not a security boundary.

To debug: k run dnsdebug --rm -it --image=busybox:1.36 --restart=Never -- nslookup kubernetes.default. A timeout means CoreDNS is down or UDP/TCP 53 is blocked; NXDOMAIN means the name is wrong. Beware dig, which ignores the search list.

Also asked: All pods in a namespace fail to resolve names. How do you debug it? · What is the Corefile, and where does it live? · What is the difference between dnsPolicy ClusterFirst and Default?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.