OnCallReady

NetworkingLinuxKubernetesSRE · 5 min read

DNS resolution failed for every service: resolv.conf, systemd-resolved, CoreDNS

When every app fails name lookups at the same minute, one shared resolver is down. Find it with resolv.conf, getent vs dig, and one query per hop.

16:07, the channel fills up: every app is failing with "unknown host". It started at 16:03. Nothing was deployed to the apps. From inside one of them:

terminal
$ kubectl exec -n shop-dns toolbox -- nslookup kubernetes.default
;; connection timed out; no servers could be reached

command terminated with exit code 1

One service failing one name is about that name or that service. Every service failing in the same minute points at what they all share: the resolver.

What a lookup goes through

An application does not speak DNS itself. It calls the C library (getaddrinfo()), and the library follows a chain of configuration files:

  1. /etc/nsswitch.conf - the order of sources. hosts: files dns means /etc/hosts first, then DNS.
  2. /etc/resolv.conf - which DNS server to ask (nameserver), and which suffixes to try for short names (search).
  3. That server, which may answer from its own cache or forward the question further upstream.

On an Ubuntu host the server in resolv.conf is not a real DNS server. It is systemd-resolved, a local caching resolver on 127.0.0.53, called the stub:

terminal
$ ls -l /etc/resolv.conf
lrwxrwxrwx 1 root root 39 Sep 14 17:43 /etc/resolv.conf -> ../run/systemd/resolve/stub-resolv.conf
$ cat /etc/resolv.conf
# This is /run/systemd/resolve/stub-resolv.conf managed by man:systemd-resolved(8).
nameserver 127.0.0.53
options edns0 trust-ad
search .
$ resolvectl status
...
Link 2 (enp0s1)
Current DNS Server: 10.64.0.1
       DNS Servers: 10.64.0.1

In a Kubernetes pod, resolv.conf points at the cluster DNS Service instead, which is CoreDNS:

terminal
$ kubectl exec -n shop-dns toolbox -- cat /etc/resolv.conf
search shop-dns.svc.cluster.local svc.cluster.local cluster.local
nameserver 10.96.0.10
options ndots:5

Either way, every process on that host or in that cluster depends on one address.

The diagnosis path

1. Read the error type: timeout, NXDOMAIN or SERVFAIL

They mean different things:

  • connection timed out; no servers could be reached - nobody answered at all. The resolver itself is down or unreachable. It is the DNS cousin of a TCP timeout, and it points the same way as one in an nginx error.log: silence, not refusal.
  • NXDOMAIN - a server answered: "this name does not exist". DNS works; the name or the search path is wrong.
  • SERVFAIL - a server answered but could not get an answer from further up the chain.

A timeout for every name, from every app, is the signature of a dead resolver.

2. Find out which resolver the client uses

cat /etc/resolv.conf where the failing app runs. A container may not use the host's resolver, and a pod almost never does.

3. Ask the way the app asks, then ask DNS directly

getent hosts NAME goes through nsswitch exactly like the application, /etc/hosts included. dig skips nsswitch and /etc/hosts and talks DNS to one server:

terminal
$ getent hosts api.lab
10.0.3.20       api.lab
$ dig +short @10.64.0.1 api.lab
10.0.3.20

Walk outwards one hop at a time: dig @127.0.0.53 (the stub), then dig @10.64.0.1 (the upstream from resolvectl status). Upstream answers, stub does not: systemd-resolved on this box is the problem (systemctl status systemd-resolved, journalctl -u systemd-resolved). If neither answers, it is the upstream or the network to it. A stopped systemd-resolved is the loud version: nothing listens on 127.0.0.53 any more, so dig prints ;; communications error to 127.0.0.53#53: connection refused three times and then ;; no servers could be reached, ping and getent say Temporary failure in name resolution, and curl says Could not resolve host.

When getent and dig disagree, the answer lives on this box: an /etc/hosts entry, or nsswitch. That is a different incident ("dig works, the app does not").

4. In a cluster: is anything behind the DNS Service?

terminal
$ kubectl get pods -n kube-system -l k8s-app=kube-dns
NAME                       READY   STATUS             RESTARTS      AGE
coredns-d6qs9mxzfp-gfn4m   0/1     CrashLoopBackOff   4 (23s ago)   102s
coredns-d6qs9mxzfp-kbw84   0/1     CrashLoopBackOff   4 (23s ago)   102s

Both replicas are crash-looping, so nothing answers on 10.96.0.10: the kube-dns Service is there, but like any Service with no ready endpoints it leads nowhere. (The label is still k8s-app=kube-dns for historical reasons; the pods run CoreDNS.) CrashLoopBackOff means start, exit, wait, retry - the reason is in the previous container's log.

5. Read why it exits

terminal
$ kubectl logs -n kube-system -l k8s-app=kube-dns --tail=3
/etc/coredns/Corefile:3 - Error during parsing: Unknown directive 'lgo'
/etc/coredns/Corefile:3 - Error during parsing: Unknown directive 'lgo'

CoreDNS reads its configuration, the Corefile, from the coredns ConfigMap:

terminal
$ kubectl get cm coredns -n kube-system -o jsonpath='{.data.Corefile}'
.:53 {
    errors
    lgo
    health {
       lameduck 5s
    }
    ...
    cahce 30 {
...

Two typos: lgo (meant log) and cahce (meant cache).

Why it broke at 16:03 and not yesterday

The config was edited yesterday, and "the logs were clean". CoreDNS's reload plugin watches the Corefile. When a new version fails to parse, the running server logs an error and keeps serving the old configuration. Nothing broke.

At 16:03 a patching job restarted the kube-system deployments. A fresh CoreDNS has no old config to fall back on: it parses the broken file, exits, and goes into CrashLoopBackOff. Both replicas read the same ConfigMap, so both died. The restart was the trigger; the typo was the cause.

The fix

Fix only the typos and keep everything else in the Corefile (forwarders, stub zones):

terminal
$ kubectl edit cm coredns -n kube-system
configmap/coredns edited
$ kubectl rollout restart deploy coredns -n kube-system
deployment.apps/coredns restarted
$ kubectl rollout status deploy coredns -n kube-system
deployment "coredns" successfully rolled out
$ kubectl exec -n shop-dns toolbox -- nslookup kubernetes.default
Server:		10.96.0.10
Address:	10.96.0.10:53

Name:	kubernetes.default.svc.cluster.local
Address: 10.96.0.1

The restart matters: crash-looping pods wait longer before each retry (up to five minutes), and a changed ConfigMap takes a while to reach running pods.

Keeping it from coming back

  • Validate resolver config before it lands. Run the new Corefile in a test pipeline or on one test instance. A hot reload that silently keeps the old config is a delayed outage.
  • Replicas do not protect against shared config. Two CoreDNS pods survive a node failure, not a bad ConfigMap. The same goes for any app that reads its settings from one: a missing ConfigMap key stops every replica together.
  • Restarting something is a test of its config on disk. After any resolver change, restart one replica on purpose and watch it come up.
  • On hosts, treat /etc/resolv.conf as managed: on Ubuntu it is a symlink owned by systemd-resolved, and DNS servers belong in netplan or resolved.conf, not in a hand edit that the next reboot or DHCP renewal replaces.

Practise it

Incident: "every service lost DNS at 16:03" (16.18) is this cluster, broken the same way. For the host side, Follow one lookup through the box (8.17) walks resolv.conf, resolved and the upstream hop by hop, and Incident: dig works, the app does not (8.23) is the getent-vs-dig case.

OnCallReady is free, with no ads and no tracking. RSS · All posts