Why DaemonSets
Some programs must run on every machine: the agent that ships each node's logs, the one that gives pods their network, the one that collects machine metrics. With a Deployment you would guess a replica count and hope the scheduler spreads them; add a node and it has no agent. A DaemonSet says "one pod on every node" and keeps it true as nodes come and go.
What you need to know already: Deployments, selectors and templates (15.16), taints on the control plane (15.7), node labels (15.7's describe node), rolling updates and rollout status/history/undo (15.16).
One per node, automatically
A DaemonSet runs one pod on every node (or every node matching a selector), adds one when a node joins, and removes it when a node leaves. The use cases are all "something every machine needs":
log shippers fluent-bit, vector - read /var/log/pods on the host
node metrics node-exporter and vendor monitoring agents
networking calico-node, cilium, kube-proxy
storage disk driver node plugins
security falco, antivirus agents
(The names are real tools; you don't need to know them - the pattern is what matters.)
apiVersion: apps/v1
kind: DaemonSet
metadata:
name: node-agent
spec:
selector:
matchLabels:
app: node-agent
template:
metadata:
labels:
app: node-agent
spec:
containers:
- name: agent
image: busybox:1.36
command: ['sh', '-c', 'while true; do echo "scrape $(hostname)"; sleep 10; done']
It is a Deployment without replicas - the node count decides. (command: sets the container's command, like the part after -- in k run. There is no create daemonset generator; a common trick is to generate a Deployment with $do and edit it.)
# after applying the manifest above (the next mission does)
k get ds node-agent
NAME DESIRED CURRENT READY UP-TO-DATE AVAILABLE NODE SELECTOR AGE
node-agent 2 2 2 2 2 <none> 6s
| column | means |
|---|---|
| DESIRED | nodes that should run a pod |
| CURRENT | pods that exist |
| READY | pods that are ready |
| UP-TO-DATE | pods on the newest template |
| AVAILABLE | pods ready long enough to count |
| NODE SELECTOR | the label rule limiting which nodes (below); <none> = all |
Why 2 and not 3
DESIRED 2 on a three-node cluster is the lesson. The DaemonSet controller respects taints (15.7): cp-1 has node-role.kubernetes.io/control-plane:NoSchedule and this pod does not tolerate it. kube-proxy and calico-node show 3:
$ k get ds -n kube-system
NAME DESIRED CURRENT READY UP-TO-DATE AVAILABLE NODE SELECTOR AGE
calico-node 3 3 3 3 3 kubernetes.io/os=linux 12d
kube-proxy 3 3 3 3 3 kubernetes.io/os=linux 12d
because their templates say:
tolerations:
- operator: Exists # tolerate EVERY taint
A toleration is the pod's "I accept this keep-out mark". To run your agent on the control plane too, tolerate just that one taint - match its key and effect:
spec:
template:
spec:
tolerations:
- key: node-role.kubernetes.io/control-plane
operator: Exists
effect: NoSchedule
For a log shipper at a bank you usually want this: logs from the control plane nodes matter as much as any.
The controller also adds some tolerations automatically, so node agents keep running on sick nodes (not ready, unreachable, under memory or disk pressure). That is why the command that empties a node for maintenance, kubectl drain, needs --ignore-daemonsets: DaemonSet pods would come straight back.
Only some nodes: nodeSelector
A nodeSelector in the pod template limits it to nodes with a given label:
spec:
template:
spec:
nodeSelector:
disk: ssd
k label node worker-1 disk=ssd puts the label disk=ssd on worker-1; k label node worker-1 disk- (key followed by -) removes it:
$ k label node worker-1 disk=ssd
node/worker-1 labeled
# ssd-tool = a DaemonSet with the nodeSelector above
k get ds
NAME DESIRED CURRENT READY UP-TO-DATE AVAILABLE NODE SELECTOR AGE
ssd-tool 1 1 1 1 1 disk=ssd 20s
$ k label node worker-1 disk-
node/worker-1 unlabeled
k get ds
NAME DESIRED CURRENT READY UP-TO-DATE AVAILABLE NODE SELECTOR AGE
ssd-tool 0 0 0 0 0 disk=ssd 40s
The DaemonSet follows the labels live: label a node and a pod appears; remove the label and it is deleted. (For ordinary pods, a nodeSelector only matters at scheduling time - removing the label later does not move them.)
How DaemonSet pods get placed
The controller creates one pod per eligible node, each with a rule pinning it to that node's name (a node affinity), and the normal scheduler binds it. You can see the rule on kube-proxy's pods (a long jsonpath: the pod's spec.affinity...values[0] is the node name):
$ k get pods -n kube-system -l k8s-app=kube-proxy -o jsonpath='{.items[0].metadata.name} {.items[0].spec.affinity.nodeAffinity.requiredDuringSchedulingIgnoredDuringExecution.nodeSelectorTerms[0].matchFields[0].values[0]}{"\n"}'
kube-proxy-r47xw worker-2
Updates
updateStrategy: RollingUpdate (default, maxUnavailable: 1) replaces the pods node by node. OnDelete waits for you to delete them. rollout status, history and undo work on DaemonSets too.
What you can now do:
- write a DaemonSet and read DESIRED/CURRENT/READY
- explain why it skips the control plane, and add the toleration to include it
- limit it to labelled nodes with a nodeSelector