Why the cluster needs its own DNS
In 16.1 a pod reached web, web.shop and web.shop.svc.cluster.local - names nobody ever typed into a DNS server. Services come and go all day; no human could keep a zone file up to date. So the cluster runs its own DNS server that reads Services straight from the API. When "nothing can reach anything", it is often this server, so you need to know how it works.
What you need to know already: how a name becomes an address, resolv.conf, search domains (8.16), dig (8.18), record types A, AAAA, CNAME, PTR, TTLs and NXDOMAIN (8.22), ndots (8.27), Services and ClusterIP (16.1), headless Services and SRV records (16.6, 16.9), ConfigMaps (15.29), CoreDNS named as a cluster add-on (15.7).
Who answers
Every pod's /etc/resolv.conf (8.16) points at one address:
$ k exec -n shop toolbox -- cat /etc/resolv.conf
search shop.svc.cluster.local svc.cluster.local cluster.local
nameserver 10.96.0.10
options ndots:5
Line by line: search = the search domains, tried in order for short names; nameserver = the DNS server to ask; options ndots:5 = the rule from 8.27 (next lesson).
The kubelet writes that file when it starts the pod. Its config holds clusterDNS: [10.96.0.10] (the server) and clusterDomain: cluster.local (the cluster domain: the DNS suffix all cluster names end in).
10.96.0.10 is the ClusterIP of Service kube-dns in kube-system. The name is historical (the older DNS server was called kube-dns); behind it run the CoreDNS pods. get svc,deploy lists both kinds in one command, and -l k8s-app=kube-dns picks them by label:
$ k get svc,deploy -n kube-system -l k8s-app=kube-dns
NAME TYPE CLUSTER-IP EXTERNAL-IP PORT(S) AGE
service/kube-dns ClusterIP 10.96.0.10 <none> 53/UDP,53/TCP,9153/TCP 13d
NAME READY UP-TO-DATE AVAILABLE AGE
deployment.apps/coredns 2/2 2 2 13d
Port 53 over UDP and TCP is DNS; 9153 is where CoreDNS publishes its own counters (queries served, errors) for monitoring tools. The Deployment runs 2 CoreDNS pods, both ready.
So DNS is just another Service: queries go through kube-proxy's rules (16.3) to one of the CoreDNS pods, NetworkPolicies (16.29) apply to them, and if CoreDNS is down, every name lookup in the cluster fails at once.
The records
CoreDNS answers these names. For Service web in namespace shop (cluster domain cluster.local):
web.shop.svc.cluster.local A the ClusterIP
_http._tcp.web.shop.svc.cluster.local SRV 0 100 80 web.shop.svc.cluster.local.
(only for NAMED ports)
headless Service: web.shop.svc.cluster.local A one record per ready pod IP
StatefulSet pod: web-0.web.shop.svc.cluster.local A that pod (hostname.subdomain.ns.svc...)
any pod: 10-244-1-16.shop.pod.cluster.local A 10.244.1.16 (the IP with dashes)
ExternalName: db.shop.svc.cluster.local CNAME the external name
reverse: 147.142.96.10.in-addr.arpa PTR web.shop.svc.cluster.local.
The pattern is always <service>.<namespace>.svc.<cluster domain>. The svc part says "this is a Service" (pods live under pod). The last line is a reverse lookup (8.22): IP to name.
nslookup NAME asks the configured server and prints the answer. The first two lines say which server answered; Name / Address is the result. Given an IP, nslookup does the reverse lookup:
$ k exec -n shop toolbox -- nslookup web
Server: 10.96.0.10
Address: 10.96.0.10:53
Name: web.shop.svc.cluster.local
Address: 10.96.142.147
$ k exec -n shop toolbox -- nslookup 10.96.142.147
Server: 10.96.0.10
Address: 10.96.0.10:53
147.142.96.10.in-addr.arpa name = web.shop.svc.cluster.local.
The pod records (10-244-1-16.shop.pod) come from the line pods insecure in CoreDNS's config (below): CoreDNS answers for any dashed IP without checking that such a pod exists. They exist so that a wildcard TLS certificate (*.shop.pod.cluster.local, 9.15) can work; almost nothing else should use them.
Names resolve across namespaces - DNS is not a security boundary:
web only from inside namespace shop (search list)
web.shop from anywhere ("web" in "shop")
web.shop.svc from anywhere
web.shop.svc.cluster.local from anywhere - the only form that is unambiguous
nslookup, dig and the search list
busybox's nslookup walks the search list and prints every miss. dig does not use the search list unless told to. nicolaka/netshoot is an image full of network tools (dig, curl, tcpdump...), handy as a test pod:
$ k run t -n shop --image=nicolaka/netshoot -- sleep 1d
pod/t created
$ k exec -n shop t -- dig +short web.shop.svc.cluster.local
10.96.142.147
$ k exec -n shop t -- dig web
;; ->>HEADER<<- opcode: QUERY, status: NXDOMAIN, id: 25838
;; QUESTION SECTION:
;web. IN A
$ k exec -n shop t -- dig +search +short web.shop
10.96.142.147
+short prints only the answer; +search makes dig use the search list. dig web asked for a top-level domain called web (like com): NXDOMAIN. That is dig being literal, not DNS being broken - a classic false alarm in incident channels.
Test with the full name, or +search, or nslookup / getent hosts NAME. getent hosts asks exactly like applications do - through the C library's resolver (libc), with the search list.
The Corefile
CoreDNS's configuration file is the Corefile. On a cluster it lives in the ConfigMap coredns in kube-system; the jsonpath prints just that key:
$ k get cm coredns -n kube-system -o jsonpath='{.data.Corefile}'
.:53 {
errors # log errors to stdout
health { # :8080/health - the liveness probe
lameduck 5s
}
ready # :8181/ready - the readiness probe
kubernetes cluster.local in-addr.arpa ip6.arpa {
pods insecure # answer the 10-244-1-16.ns.pod records
fallthrough in-addr.arpa ip6.arpa
ttl 30
}
prometheus :9153 # metrics
forward . /etc/resolv.conf { # everything else: the node's upstream
max_concurrent 1000
}
cache 30 {
disable success cluster.local
disable denial cluster.local
}
loop # detect forwarding loops and exit
reload # re-read this file when it changes
loadbalance # shuffle A records
}
The file is one server block: .:53 { ... } = "for every name (. is the DNS root, so everything), on port 53". Inside, one plugin per line - each plugin is one feature. Order in the file does not matter: CoreDNS runs plugins in a fixed order built into it.
What each one does:
errors print errors to the pod's log
health an HTTP page on :8080/health the kubelet checks to see CoreDNS is
alive; lameduck 5s = keep answering 5 s after being told to stop
ready an HTTP page on :8181/ready: "I have loaded everything"
kubernetes answer the cluster zone from the API (below)
(metrics) the :9153 line publishes CoreDNS's counters for monitoring tools
forward send every other name to an upstream DNS server
cache 30 remember answers for up to 30 s (but not for cluster.local here)
loop detect a forwarding loop and stop
reload re-read the Corefile when it changes
loadbalance shuffle the order of A records in each answer
The important ones in detail:
- kubernetes: the zone
cluster.local(and reverse lookups,in-addr.arpa) is answered from the API - Services and EndpointSlices - not from files.fallthrough= if the name is not found here, let the next plugin try. - forward . /etc/resolv.conf: anything else goes to the servers in that file. CoreDNS runs with
dnsPolicy: Default(below), so its /etc/resolv.conf is the node's. On these Ubuntu nodes the kubelet hands it/run/systemd/resolve/resolv.conf(nameserver 10.64.0.1), not the127.0.0.53local stub of 8.16. - Why that matters: if CoreDNS's resolv.conf said 127.0.0.53, inside the pod that is CoreDNS itself - it would forward to itself forever.
loopdetects that and CoreDNS exits with "Loop ... detected": a famous CrashLoopBackOff (15.14) on fresh Ubuntu clusters. - reload: CoreDNS checks the Corefile every 30 s (plus a random extra) and applies a changed one. Add the kubelet's delay in updating ConfigMap files inside pods (15.29) and an edit takes up to about a minute and a half to land. A broken edit is logged and ignored until the pods restart.
dnsPolicy and dnsConfig
A pod's dnsPolicy decides what goes into its resolv.conf:
ClusterFirst (default) cluster DNS, search list, ndots:5
Default the node's resolver - NOT the default, despite the name
ClusterFirstWithHostNet what hostNetwork pods need to use cluster DNS
None nothing; you supply everything in dnsConfig
(A hostNetwork pod uses the node's network directly instead of its own, like Docker's --network host, 11.15.)
dnsConfig adds to or overrides the generated file; hostAliases adds lines to the pod's /etc/hosts:
spec:
dnsPolicy: ClusterFirst
dnsConfig:
options:
- name: ndots
value: "2"
searches:
- corp.lab
hostAliases: # extra /etc/hosts lines, per pod
- ip: 10.0.3.12
hostnames: [legacy-db]
dnsConfig is merged into what the policy generates: nameservers appended (max 3), searches appended, options overridden by name. With dnsPolicy: None and no dnsConfig.nameservers, the pod has no nameserver at all.
Debugging DNS from a pod
k run dnsdebug --rm -it --image=busybox:1.36 --restart=Never -- nslookup kubernetes.default
k exec -it deploy/api -- cat /etc/resolv.conf
k get pods -n kube-system -l k8s-app=kube-dns # running? ready?
k get endpointslice -n kube-system -l kubernetes.io/service-name=kube-dns
k logs -n kube-system -l k8s-app=kube-dns --tail=20
Line by line:
run ... --rm -it --restart=Never a one-off pod: interactive, never restarted,
deleted when the command ends
exec deploy/api exec into any one pod of Deployment api
get pods -l k8s-app=kube-dns are the CoreDNS pods running and Ready?
get endpointslice ... does the kube-dns Service have endpoints?
logs ... --tail=20 the last 20 lines from every CoreDNS pod
nslookup kubernetes.default is the standard test (kubernetes is the Service for the API server, in namespace default): it exercises the search list, CoreDNS, the kubernetes plugin and the API.
- If it fails with
;; connection timed out; no servers could be reached, the query never got an answer: CoreDNS down, or a NetworkPolicy dropping UDP 53. - If it returns NXDOMAIN, DNS works and the name is wrong.
What you can now do:
- build the DNS name of any Service or pod, from any namespace
- read the Corefile and say what each plugin does
- test DNS from a pod without being fooled by dig's missing search list