OnCallReady

Lesson 17.26 · Kubernetes: Scheduling, Health & Security · 12 min read

PodDisruptionBudgets and why they matter during a drain

In plain words

Imagine a hospital ward that must always have at least two nurses on duty. Nurses can still get sick without warning (nobody can stop that). But when the manager wants to send nurses on a training course, the rule is checked first: "you can send one away only if two are still left." If only two are working, the manager has to wait.

A PodDisruptionBudget is that rule for voluntary disruptions: kubectl drain, cluster autoscaler scale-downs, anything using the Eviction API. The API server refuses an eviction that would drop healthy pods below minAvailable (or above maxUnavailable). It can't help with crashes, node failures or OOM kills, and kubectl delete pod ignores it entirely.

Voluntary vs involuntary

The problem. Every node needs maintenance: kernel patches, Kubernetes upgrades. Emptying a node (a drain) moves its pods elsewhere - and if two replicas of the same service happen to be moved at once, the service is down. A budget per service tells the drain how many pods it may take away at a time.

What you need to know already: eviction (17.3), cordon and taints (17.13), Deployments and ReplicaSets (15.16), DaemonSets (15.22), emptyDir volumes (16.37).

A PodDisruptionBudget (PDB) limits how many pods of a set may be down because somebody chose to take them down:

voluntary (a PDB constrains these)involuntary (a PDB cannot help)
kubectl drain for node maintenance / upgradesnode hardware failure, kernel panic
cluster autoscaler removing an underused nodethe VM is deleted by the cloud
a person or tool evicting pods through the Eviction APInode-pressure eviction by the kubelet
OOM kills, crashes

The mechanism is the Eviction API (POST .../pods/NAME/eviction): instead of deleting a pod, a client asks for it to be evicted, and the apiserver refuses if that would break a matching PDB. kubectl drain uses it; kubectl delete pod does not - deleting a pod ignores PDBs completely.

Writing one

apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
  name: web
  namespace: shop
spec:
  minAvailable: 2          # OR maxUnavailable: 1 - never both
  selector:
    matchLabels: {app: web}

kubectl create pdb NAME --selector=LABEL --max-unavailable=N (or --min-available=N) creates one from the command line:

# an illustration: a PDB on shop/web (the PDB mission)
kubectl create pdb web -n shop --selector=app=web --max-unavailable=1
poddisruptionbudget.policy/web created
kubectl get pdb -n shop
NAME   MIN AVAILABLE   MAX UNAVAILABLE   ALLOWED DISRUPTIONS   AGE
web    N/A             1                 1                     5s
kubectl describe pdb web -n shop
Name:           web
Namespace:      shop
Max unavailable:  1
Selector:       app=web
Status:
    Allowed disruptions:  1
    Current:              3
    Desired:              2
    Total:                3

How a PDB blocks a drain forever

# one replica
spec:
  replicas: 1
---
# "we must always have one available"
spec:
  minAvailable: 1

Allowed disruptions = 1 - 1 = 0, permanently. Evicting the only pod would break the budget, so the apiserver refuses - every 5 seconds, forever. kubectl drain NODE cordons the node, then evicts its pods one by one; --ignore-daemonsets = skip DaemonSet pods (they would come straight back), --delete-emptydir-data = allow evicting pods whose emptyDir data will be lost:

# an illustration: one replica, minAvailable 1 (the PDB mission)
$ kubectl drain worker-1 --ignore-daemonsets --delete-emptydir-data
node/worker-1 cordoned
Warning: ignoring DaemonSet-managed Pods: kube-system/calico-node-2qs6n, kube-system/kube-proxy-v7zs8
evicting pod shop/web-5d8f7c9b4d-2kq9x
error when evicting pods/"web-5d8f7c9b4d-2kq9x" -n "shop" (will retry after 5s): Cannot evict pod as it would violate the pod's disruption budget.
evicting pod shop/web-5d8f7c9b4d-2kq9x
error when evicting pods/"web-5d8f7c9b4d-2kq9x" -n "shop" (will retry after 5s): Cannot evict pod as it would violate the pod's disruption budget.
...

The node upgrade hangs, the maintenance window runs out, and on a managed cloud cluster the automatic node upgrade fails with a PDB error. "Very common self-inflicted wound." The fixes, best first: run at least 2 replicas (then minAvailable 1 allows one disruption); use maxUnavailable: 1 instead; or, for a true singleton, decide consciously that it may be disrupted and do not give it a PDB. As a last resort the person draining can use kubectl drain --disable-eviction (delete instead of evict - ignores the PDB) or --timeout so the command gives up instead of hanging.

A drain also fails fast for things it will not touch without permission:

error: unable to drain node "worker-1" due to error: [cannot delete DaemonSet-managed Pods (use --ignore-daemonsets to ignore): kube-system/calico-node-7xk2p, cannot delete Pods with local storage (use --delete-emptydir-data to override): shop/cache-6f4d8-9zt2m, cannot delete Pods that declare no controller (use --force to override): default/debug], continuing command...

--force deletes bare pods for good - nothing recreates them.

Unhealthy pods and the budget

What if the pods covered by the PDB are already broken (CrashLoopBackOff, never Ready)? With the default unhealthyPodEvictionPolicy: IfHealthyBudget, a not-Ready pod may only be evicted when the budget is currently satisfied - so a Deployment that is fully broken can still block a drain. unhealthyPodEvictionPolicy: AlwaysAllow lets not-Ready pods be evicted regardless (they are not serving anyway) and is the setting most platform teams recommend.

The drain in practice

# an illustration: a PDB on shop/web (the PDB mission)
$ kubectl drain worker-1 --ignore-daemonsets --delete-emptydir-data --timeout=10m
... maintenance ...
$ kubectl uncordon worker-1

Drain = cordon + evict everything evictable, respecting PDBs, waiting for each pod to terminate (its grace period, preStop hooks - 3.18). Evicted pods are recreated by their controllers elsewhere - which requires capacity elsewhere. And after uncordon, nothing moves back by itself: the node stays empty until new pods are scheduled.

What you can now do

Why it helps

PDBs are what make node maintenance safe: without them, draining a node during patching can take down every replica of a service that happened to be there. With them wrong, the opposite happens, and it's the famous self-inflicted wound: a one-replica Deployment with minAvailable: 1 blocks kubectl drain forever, and on a managed cloud cluster the automatic node upgrade fails with a PDB error in the middle of the maintenance window.

On a platform team you'll set PDB policy for tenants (maxUnavailable: 1, unhealthyPodEvictionPolicy: AlwaysAllow), debug stuck drains during cluster upgrades, and explain why a broken Deployment can still block maintenance. It's also a regular CKA and interview topic, usually alongside drains.

Commands in this lesson

kubectl

FAQ

Does a PDB protect against node failures?

No. It only constrains voluntary disruptions that go through the Eviction API: drains, autoscaler scale-downs, operators evicting pods. Hardware failures, kernel panics, deleted VMs, node-pressure evictions by the kubelet and OOM kills happen regardless. Protection against those comes from replicas spread across nodes and zones.

Why is my drain stuck with "Cannot evict pod as it would violate the pod's disruption budget"?

Allowed disruptions is 0. The classic case: one replica with minAvailable: 1, so evicting the only pod would always break the budget, and drain retries every 5 seconds forever. It also happens when the covered pods are already unhealthy. Fix it with more replicas, maxUnavailable: 1, or no PDB for true singletons.

minAvailable or maxUnavailable?

Prefer maxUnavailable. It keeps working as you scale: maxUnavailable: 1 always allows one disruption as long as all pods are healthy, whether you run 2 replicas or 20. minAvailable with a fixed number can become 0 allowed disruptions when someone scales down. You can't set both on one PDB.

Does kubectl delete pod respect the PDB?

No. Deleting a pod skips the Eviction API, so PDBs are ignored completely. kubectl drain uses eviction and respects them, unless you pass --disable-eviction, which switches it to deletes. That's a last resort when a PDB blocks maintenance and you've decided the disruption is acceptable.

What does unhealthyPodEvictionPolicy: AlwaysAllow change?

With the default IfHealthyBudget, a pod that isn't Ready may only be evicted when the budget is currently satisfied, so a fully broken Deployment can still block a drain. AlwaysAllow lets not-Ready pods be evicted regardless, since they aren't serving anyway. Most platform teams recommend it.

In an interview Junior

What is a PodDisruptionBudget?

A PDB limits how many pods of a set may be down because of a voluntary disruption - a kubectl drain, a node upgrade, the cluster autoscaler removing a node. It cannot help against crashes, node failures or kubelet evictions.

apiVersion: policy/v1
kind: PodDisruptionBudget
spec:
  maxUnavailable: 1
  selector: {matchLabels: {app: web}}

It works through the Eviction API: kubectl drain cordons the node and asks to evict each pod; the API server refuses when that would break the budget, and the drain retries. kubectl delete pod ignores PDBs entirely. kubectl get pdb shows ALLOWED DISRUPTIONS.

The classic self-inflicted wound: one replica with minAvailable: 1 - allowed disruptions is 0 forever, and every drain or node upgrade hangs. Run at least two replicas, prefer maxUnavailable, and consider unhealthyPodEvictionPolicy: AlwaysAllow so broken pods do not block drains.

Also asked: A node pool upgrade is stuck. How do you find and fix the cause? · What flags does kubectl drain usually need, and why? · What is the difference between a voluntary and an involuntary disruption?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.