OnCallReady

Lesson 15.22 · Kubernetes: Architecture & Workloads · 15 min read

DaemonSets: one pod per node, and why the control plane is left out

In plain words

Every classroom in a school needs a fire extinguisher. You don't order "20 fire extinguishers" and hope they end up in the right places; you say "one in every classroom". Build a new classroom and it gets one automatically; close a classroom and its extinguisher goes. Classrooms marked "staff only" don't get one unless the rule says it's allowed there.

A DaemonSet is that rule for pods: one pod on every node, or every node matching a nodeSelector, with no replicas field. It's used for log shippers, node-exporter, CNI agents like calico-node, kube-proxy and CSI node plugins. It respects taints, which is why it shows DESIRED 2 on a 3-node cluster unless its pods tolerate the control-plane taint. It follows node labels live, and updates node by node.

Why DaemonSets

Some programs must run on every machine: the agent that ships each node's logs, the one that gives pods their network, the one that collects machine metrics. With a Deployment you would guess a replica count and hope the scheduler spreads them; add a node and it has no agent. A DaemonSet says "one pod on every node" and keeps it true as nodes come and go.

What you need to know already: Deployments, selectors and templates (15.16), taints on the control plane (15.7), node labels (15.7's describe node), rolling updates and rollout status/history/undo (15.16).

One per node, automatically

A DaemonSet runs one pod on every node (or every node matching a selector), adds one when a node joins, and removes it when a node leaves. The use cases are all "something every machine needs":

log shippers        fluent-bit, vector - read /var/log/pods on the host
node metrics        node-exporter and vendor monitoring agents
networking          calico-node, cilium, kube-proxy
storage             disk driver node plugins
security            falco, antivirus agents

(The names are real tools; you don't need to know them - the pattern is what matters.)

apiVersion: apps/v1
kind: DaemonSet
metadata:
  name: node-agent
spec:
  selector:
    matchLabels:
      app: node-agent
  template:
    metadata:
      labels:
        app: node-agent
    spec:
      containers:
      - name: agent
        image: busybox:1.36
        command: ['sh', '-c', 'while true; do echo "scrape $(hostname)"; sleep 10; done']

It is a Deployment without replicas - the node count decides. (command: sets the container's command, like the part after -- in k run. There is no create daemonset generator; a common trick is to generate a Deployment with $do and edit it.)

# after applying the manifest above (the next mission does)
k get ds node-agent
NAME         DESIRED   CURRENT   READY   UP-TO-DATE   AVAILABLE   NODE SELECTOR   AGE
node-agent   2         2         2       2            2           <none>          6s
columnmeans
DESIREDnodes that should run a pod
CURRENTpods that exist
READYpods that are ready
UP-TO-DATEpods on the newest template
AVAILABLEpods ready long enough to count
NODE SELECTORthe label rule limiting which nodes (below); <none> = all

Why 2 and not 3

DESIRED 2 on a three-node cluster is the lesson. The DaemonSet controller respects taints (15.7): cp-1 has node-role.kubernetes.io/control-plane:NoSchedule and this pod does not tolerate it. kube-proxy and calico-node show 3:

$ k get ds -n kube-system
NAME          DESIRED   CURRENT   READY   UP-TO-DATE   AVAILABLE   NODE SELECTOR            AGE
calico-node   3         3         3       3            3           kubernetes.io/os=linux   12d
kube-proxy    3         3         3       3            3           kubernetes.io/os=linux   12d

because their templates say:

tolerations:
- operator: Exists          # tolerate EVERY taint

A toleration is the pod's "I accept this keep-out mark". To run your agent on the control plane too, tolerate just that one taint - match its key and effect:

spec:
  template:
    spec:
      tolerations:
      - key: node-role.kubernetes.io/control-plane
        operator: Exists
        effect: NoSchedule

For a log shipper at a bank you usually want this: logs from the control plane nodes matter as much as any.

The controller also adds some tolerations automatically, so node agents keep running on sick nodes (not ready, unreachable, under memory or disk pressure). That is why the command that empties a node for maintenance, kubectl drain, needs --ignore-daemonsets: DaemonSet pods would come straight back.

Only some nodes: nodeSelector

A nodeSelector in the pod template limits it to nodes with a given label:

spec:
  template:
    spec:
      nodeSelector:
        disk: ssd

k label node worker-1 disk=ssd puts the label disk=ssd on worker-1; k label node worker-1 disk- (key followed by -) removes it:

$ k label node worker-1 disk=ssd
node/worker-1 labeled
# ssd-tool = a DaemonSet with the nodeSelector above
k get ds
NAME       DESIRED   CURRENT   READY   UP-TO-DATE   AVAILABLE   NODE SELECTOR   AGE
ssd-tool   1         1         1       1            1           disk=ssd        20s
$ k label node worker-1 disk-
node/worker-1 unlabeled
k get ds
NAME       DESIRED   CURRENT   READY   UP-TO-DATE   AVAILABLE   NODE SELECTOR   AGE
ssd-tool   0         0         0       0            0           disk=ssd        40s

The DaemonSet follows the labels live: label a node and a pod appears; remove the label and it is deleted. (For ordinary pods, a nodeSelector only matters at scheduling time - removing the label later does not move them.)

How DaemonSet pods get placed

The controller creates one pod per eligible node, each with a rule pinning it to that node's name (a node affinity), and the normal scheduler binds it. You can see the rule on kube-proxy's pods (a long jsonpath: the pod's spec.affinity...values[0] is the node name):

$ k get pods -n kube-system -l k8s-app=kube-proxy -o jsonpath='{.items[0].metadata.name} {.items[0].spec.affinity.nodeAffinity.requiredDuringSchedulingIgnoredDuringExecution.nodeSelectorTerms[0].matchFields[0].values[0]}{"\n"}'
kube-proxy-r47xw worker-2

Updates

updateStrategy: RollingUpdate (default, maxUnavailable: 1) replaces the pods node by node. OnDelete waits for you to delete them. rollout status, history and undo work on DaemonSets too.

What you can now do:

Why it helps

Every platform runs DaemonSets for logging, monitoring and security agents, and you'll write and debug them. Situations: control-plane logs are missing from the log platform because the log shipper doesn't tolerate the control-plane taint; a node-exporter isn't running on new GPU nodes because they have a custom taint; kubectl drain fails until you add --ignore-daemonsets, because DaemonSet pods would come straight back. You'll also meet the adoption trap: a test pod with the same labels as a DaemonSet's gets adopted and one pod deleted. These come up in exam tasks and in reviews of agent deployments.

FAQ

Why does my DaemonSet show DESIRED 2 on a 3-node cluster?

The DaemonSet controller respects taints. The control-plane node has node-role.kubernetes.io/control-plane:NoSchedule, and your pods don't tolerate it, so that node isn't eligible. Add a toleration for that key with operator: Exists and effect: NoSchedule. kube-proxy and calico-node run everywhere because they tolerate every taint (operator: Exists with no key).

Why does kubectl drain need --ignore-daemonsets?

DaemonSet pods are meant to run on every node, and the controller adds tolerations like unschedulable so they stay even on cordoned nodes. Evicting them would be pointless, since they'd come straight back. --ignore-daemonsets tells drain to leave them and continue evicting everything else.

How does a DaemonSet run only on some nodes?

With nodeSelector or node affinity in the pod template. The DaemonSet follows labels live: label a node disk=ssd and a pod appears there; remove the label and the pod is deleted. That's different from node affinity on normal pods, which is "IgnoredDuringExecution" and doesn't evict running pods.

Who schedules DaemonSet pods?

The DaemonSet controller creates one pod per eligible node with a required node affinity on that node's name, and the default scheduler binds it. So scheduling features like priority and preemption apply, and the pod's affinity shows which node it was created for.

How are DaemonSets updated?

updateStrategy: RollingUpdate (the default, maxUnavailable: 1) replaces pods node by node; OnDelete updates a node's pod only when you delete it. kubectl rollout status, history and undo work on DaemonSets too. For agents on large clusters, raise maxUnavailable carefully, since every node briefly loses its agent during its update.

In an interview Junior

What is a DaemonSet, and when would you use one?

A DaemonSet runs one pod on every node (or every node matching a nodeSelector), adds one when a node joins and removes it when a node leaves. There is no replicas: the node count decides.

Use it for things every machine needs: log shippers reading /var/log/pods, node metrics agents, the network plugin (calico-node), kube-proxy.

The detail to know: it respects taints. On kubeadm the control plane has node-role.kubernetes.io/control-plane:NoSchedule, so a plain DaemonSet shows DESIRED 2 on a three-node cluster. To include the control plane, add a matching toleration to the pod template (kube-proxy tolerates everything with operator: Exists).

Updates roll node by node (maxUnavailable: 1), and rollout status/history/undo work as for Deployments. It is also why kubectl drain needs --ignore-daemonsets.

Also asked: Why does a DaemonSet not run on the control plane node by default? · How do you run a DaemonSet only on some nodes? · What is the difference between a taint and a toleration?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.