Voluntary vs involuntary
The problem. Every node needs maintenance: kernel patches, Kubernetes upgrades. Emptying a node (a drain) moves its pods elsewhere - and if two replicas of the same service happen to be moved at once, the service is down. A budget per service tells the drain how many pods it may take away at a time.
What you need to know already: eviction (17.3), cordon and taints (17.13), Deployments and ReplicaSets (15.16), DaemonSets (15.22), emptyDir volumes (16.37).
A PodDisruptionBudget (PDB) limits how many pods of a set may be down because somebody chose to take them down:
| voluntary (a PDB constrains these) | involuntary (a PDB cannot help) |
|---|---|
kubectl drain for node maintenance / upgrades | node hardware failure, kernel panic |
| cluster autoscaler removing an underused node | the VM is deleted by the cloud |
| a person or tool evicting pods through the Eviction API | node-pressure eviction by the kubelet |
| OOM kills, crashes |
The mechanism is the Eviction API (POST .../pods/NAME/eviction): instead of deleting a pod, a client asks for it to be evicted, and the apiserver refuses if that would break a matching PDB. kubectl drain uses it; kubectl delete pod does not - deleting a pod ignores PDBs completely.
Writing one
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: web
namespace: shop
spec:
minAvailable: 2 # OR maxUnavailable: 1 - never both
selector:
matchLabels: {app: web}
kubectl create pdb NAME --selector=LABEL --max-unavailable=N (or --min-available=N) creates one from the command line:
# an illustration: a PDB on shop/web (the PDB mission)
kubectl create pdb web -n shop --selector=app=web --max-unavailable=1
poddisruptionbudget.policy/web created
kubectl get pdb -n shop
NAME MIN AVAILABLE MAX UNAVAILABLE ALLOWED DISRUPTIONS AGE
web N/A 1 1 5s
kubectl describe pdb web -n shop
Name: web
Namespace: shop
Max unavailable: 1
Selector: app=web
Status:
Allowed disruptions: 1
Current: 3
Desired: 2
Total: 3
- Current = pods matching the selector that are Ready ("healthy").
- Desired = how many must stay healthy (minAvailable, or total - maxUnavailable).
- Allowed disruptions = current - desired. 0 means no eviction will be allowed.
- Percentages are allowed (
minAvailable: 50%, rounded up). - Prefer
maxUnavailable: it keeps working when you scale the Deployment up or down.
How a PDB blocks a drain forever
# one replica
spec:
replicas: 1
---
# "we must always have one available"
spec:
minAvailable: 1
Allowed disruptions = 1 - 1 = 0, permanently. Evicting the only pod would break the budget, so the apiserver refuses - every 5 seconds, forever. kubectl drain NODE cordons the node, then evicts its pods one by one; --ignore-daemonsets = skip DaemonSet pods (they would come straight back), --delete-emptydir-data = allow evicting pods whose emptyDir data will be lost:
# an illustration: one replica, minAvailable 1 (the PDB mission)
$ kubectl drain worker-1 --ignore-daemonsets --delete-emptydir-data
node/worker-1 cordoned
Warning: ignoring DaemonSet-managed Pods: kube-system/calico-node-2qs6n, kube-system/kube-proxy-v7zs8
evicting pod shop/web-5d8f7c9b4d-2kq9x
error when evicting pods/"web-5d8f7c9b4d-2kq9x" -n "shop" (will retry after 5s): Cannot evict pod as it would violate the pod's disruption budget.
evicting pod shop/web-5d8f7c9b4d-2kq9x
error when evicting pods/"web-5d8f7c9b4d-2kq9x" -n "shop" (will retry after 5s): Cannot evict pod as it would violate the pod's disruption budget.
...
The node upgrade hangs, the maintenance window runs out, and on a managed cloud cluster the automatic node upgrade fails with a PDB error. "Very common self-inflicted wound." The fixes, best first: run at least 2 replicas (then minAvailable 1 allows one disruption); use maxUnavailable: 1 instead; or, for a true singleton, decide consciously that it may be disrupted and do not give it a PDB. As a last resort the person draining can use kubectl drain --disable-eviction (delete instead of evict - ignores the PDB) or --timeout so the command gives up instead of hanging.
A drain also fails fast for things it will not touch without permission:
error: unable to drain node "worker-1" due to error: [cannot delete DaemonSet-managed Pods (use --ignore-daemonsets to ignore): kube-system/calico-node-7xk2p, cannot delete Pods with local storage (use --delete-emptydir-data to override): shop/cache-6f4d8-9zt2m, cannot delete Pods that declare no controller (use --force to override): default/debug], continuing command...
--force deletes bare pods for good - nothing recreates them.
Unhealthy pods and the budget
What if the pods covered by the PDB are already broken (CrashLoopBackOff, never Ready)? With the default unhealthyPodEvictionPolicy: IfHealthyBudget, a not-Ready pod may only be evicted when the budget is currently satisfied - so a Deployment that is fully broken can still block a drain. unhealthyPodEvictionPolicy: AlwaysAllow lets not-Ready pods be evicted regardless (they are not serving anyway) and is the setting most platform teams recommend.
The drain in practice
# an illustration: a PDB on shop/web (the PDB mission)
$ kubectl drain worker-1 --ignore-daemonsets --delete-emptydir-data --timeout=10m
... maintenance ...
$ kubectl uncordon worker-1
Drain = cordon + evict everything evictable, respecting PDBs, waiting for each pod to terminate (its grace period, preStop hooks - 3.18). Evicted pods are recreated by their controllers elsewhere - which requires capacity elsewhere. And after uncordon, nothing moves back by itself: the node stays empty until new pods are scheduled.
What you can now do
- Write a PDB (
maxUnavailablepreferred) and read ALLOWED DISRUPTIONS. - Drain and uncordon a node, and know which flags a drain needs and why.
- Spot the "one replica + minAvailable: 1" trap that blocks every drain.