OnCallReady

Lesson 17.24 · Kubernetes: Scheduling, Health & Security · 11 min read

Rolling updates: maxSurge, maxUnavailable, and the rollout verbs

In plain words

Imagine replacing all the light bulbs in a hallway without ever leaving it dark. You could screw in one new bulb first, check it works, then take one old one out, and repeat (never less light than before). Or, if you have no spare socket, take one old one out first, then put in a new one (a bit dimmer for a moment). If a new bulb doesn't light up, you stop and keep the old ones burning.

That's a rolling update. maxSurge is how many extra sockets you may use above replicas; maxUnavailable is how many may be dark. "Lights up" means the new pod is Ready (for minReadySeconds). kubectl rollout status, history and undo are how you watch, inspect and reverse it, and Recreate is turning everything off first.

How a rolling update moves

The problem. Shipping a new version means replacing running pods. Done carelessly you get downtime (all old pods gone before new ones work) or a broken release everywhere in one minute. The rolling-update settings and the kubectl rollout commands let you ship safely and go back fast.

What you need to know already: Deployments, ReplicaSets and the basic rollout maths (15.16), readiness probes (17.20), annotations (15.26).

A Deployment's strategy decides how the old ReplicaSet is replaced by the new one:

spec:
  replicas: 4
  minReadySeconds: 10           # a new pod must stay Ready 10s before it counts as available
  progressDeadlineSeconds: 600  # no progress for 10 min -> Progressing=False
  strategy:
    type: RollingUpdate         # or Recreate: kill all old, then start new (downtime)
    rollingUpdate:
      maxSurge: 1               # at most replicas + 1 pods in total
      maxUnavailable: 0         # never fewer than replicas available

maxSurge: 1, maxUnavailable: 0 is the zero-downtime setting: start one new pod, wait until it is available (Ready for minReadySeconds), remove one old, repeat. maxSurge: 0, maxUnavailable: 1 is the "no spare capacity" setting (quota-bound namespaces, anti-affinity per node): remove first, then add.

Readiness is what makes this safe. "Available" means Ready. A new version whose readiness probe never passes stalls the rollout with the old pods still serving:

# an illustration: shop/web with the strategy above (the rollout missions)
kubectl rollout status deploy/web -n shop --timeout=60s
Waiting for deployment "web" rollout to finish: 1 out of 4 new replicas have been updated...
error: timed out waiting for the condition
kubectl get pods -n shop -l app=web
NAME                   READY   STATUS    RESTARTS   AGE
web-5d8f7c9b4d-2kq9x   1/1     Running   0          2d
web-5d8f7c9b4d-7mz4p   1/1     Running   0          2d
web-5d8f7c9b4d-q8w2x   1/1     Running   0          2d
web-5d8f7c9b4d-x9c2m   1/1     Running   0          2d
web-7c9f4d8b6c-4kq2x   0/1     Running   0          58s      <- the new version, never Ready

Without a readiness probe the new pod is "available" the moment the process starts, and a broken release replaces every good pod in a minute.

After progressDeadlineSeconds without progress the Deployment says so:

# an illustration: shop/web with the strategy above (the rollout missions)
kubectl describe deploy web -n shop | grep -A3 Conditions
Conditions:
  Type           Status  Reason
  ----           ------  ------
  Available      True    MinimumReplicasAvailable
  Progressing    False   ProgressDeadlineExceeded
kubectl rollout status deploy/web -n shop
error: deployment "web" exceeded its progress deadline

Kubernetes does not roll back automatically. ProgressDeadlineExceeded is a signal for your deployment automation (a script that runs kubectl rollout status and reacts to its exit code) or for you.

The rollout verbs

# an illustration: shop/web with the strategy above (the rollout missions)
kubectl rollout status deploy/web -n shop            # block until done or failed (deploy scripts use this)
kubectl rollout history deploy/web -n shop
deployment.apps/web
REVISION  CHANGE-CAUSE
1         <none>
2         bump to 1.28
3         bump to 1.29
kubectl rollout history deploy/web -n shop --revision=2   # the pod template of revision 2
kubectl rollout undo deploy/web -n shop                    # back to the previous revision
deployment.apps/web rolled back
kubectl rollout undo deploy/web -n shop --to-revision=1
kubectl rollout restart deploy/web -n shop                 # new pods, same spec (annotation bump)
kubectl rollout pause deploy/web -n shop                   # batch several changes...
kubectl rollout resume deploy/web -n shop                  # ...into one rollout

Later (Ch 26): GitOps tools re-apply Git continuously, which makes "revert in Git" the only real rollback.

Recreate

strategy: {type: Recreate} scales the old ReplicaSet to 0, waits for its pods to be gone, then creates the new ones. Downtime by design. Used for workloads that must not run two versions at once (a singleton - a program that must only ever run once - holding a lock, an app with an incompatible database schema migration, an RWO volume that only one pod can mount, 16.41).

What you can now do

Why it helps

Every deploy you ship is a rollout, so you'll read these fields constantly. maxSurge: 1, maxUnavailable: 0 is the zero-downtime setting; knowing how percentages round (surge up, unavailable down) explains why a 3-replica Deployment behaves differently from a 4-replica one.

In incidents, kubectl rollout undo is the fastest mitigation for a bad release, and knowing what it doesn't undo (ConfigMaps, replica counts) prevents a false sense of safety. ProgressDeadlineExceeded is the signal your deploy automation should react to, since Kubernetes never rolls back by itself. And if your YAML lives in Git and is re-applied automatically, an in-cluster undo is reverted by the next apply, so the real fix is in Git.

FAQ

What do maxSurge and maxUnavailable mean exactly?

maxSurge is how many pods above replicas may exist during the update; percentages round up. maxUnavailable is how many below replicas may be unavailable; percentages round down. Defaults are 25% and 25%, and both 0 is invalid. With 4 replicas the defaults give 1 and 1; with 3 replicas, 1 surge and 0 unavailable.

Does Kubernetes roll back automatically when a rollout fails?

No. After progressDeadlineSeconds (600 by default) without progress, the Deployment's Progressing condition becomes False with reason ProgressDeadlineExceeded, and kubectl rollout status exits with an error. That's a signal. Rolling back is up to you or your deploy script (it can react to the exit code of kubectl rollout status).

What does kubectl rollout undo actually restore?

Only the pod template of an earlier ReplicaSet. Replica count, strategy, and anything the pods read from outside (ConfigMaps, Secrets, external config) aren't part of a revision. Rolling back the image doesn't roll back a config change made in a ConfigMap. The restored revision becomes the newest number in rollout history.

Where does CHANGE-CAUSE come from?

From the kubernetes.io/change-cause annotation on the Deployment at the time of the change. --record is deprecated, so set it yourself: kubectl annotate deploy/web kubernetes.io/change-cause="bump to 1.29". Many deploy scripts set it automatically to the commit or release version.

When should I use Recreate instead of RollingUpdate?

When two versions must never run at the same time: a singleton holding a lock, an incompatible schema migration, or a single-replica app with an RWO volume that only one node can mount. Recreate scales the old ReplicaSet to zero, waits until its pods are gone, then starts the new ones. That means downtime, by design.

In an interview Junior

How does a Deployment rolling update work, and how do you make it safe?

A template change creates a new ReplicaSet; the Deployment scales it up and the old one down within two limits:

maxSurge: 1, maxUnavailable: 0 is zero downtime: one new pod, wait until it is available (Ready for minReadySeconds), remove one old, repeat.

Readiness is what makes it safe: a new version whose readiness never passes stalls the rollout with the old pods still serving. Without a readiness probe, a broken release replaces every pod within a minute. After progressDeadlineSeconds the Deployment reports ProgressDeadlineExceeded - but does not roll back on its own.

The verbs: kubectl rollout status --timeout (non-zero exit on failure - deploy scripts use it), history (with a kubernetes.io/change-cause annotation), undo --to-revision=N, restart, pause/resume. Undo only restores the pod template, not a ConfigMap you changed - and if git re-applies, revert there too.

Also asked: A release went out and errors spiked. What do you do in the first five minutes? · When would you use the Recreate strategy? · What does kubectl rollout undo restore, and what does it not?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.