OnCallReady

Lesson 15.16 · Kubernetes: Architecture & Workloads · 27 min read

Deployments and ReplicaSets: rolling updates, the maths, history and rollback

In plain words

Imagine replacing the chairs in a busy classroom while lessons continue. You bring in one new chair, a child moves onto it, then you take away one old chair, and repeat, so there's never a moment with too few seats. You keep the old chairs in the storeroom for a while, in case the new ones turn out to be wobbly and you need to swap back.

A Deployment does that for pods. It manages ReplicaSets, one per version of the pod template. A rolling update scales the new one up and the old one down, limited by maxSurge (extra pods allowed, rounds up) and maxUnavailable (pods allowed missing, rounds down). Old ReplicaSets stay at 0 replicas as history, so kubectl rollout undo can switch back, and rollout status tells a pipeline whether it worked.

Why Deployments

A bare pod dies with its node and never comes back (15.14). Real services need N copies that are kept alive, and a way to ship a new version without dropping traffic - and to go back when the new version is bad. A Deployment is that: "keep N identical pods running, and when I change them, replace them a few at a time". It is the object you will write and review most in your career.

What you need to know already: pods, READY and CrashLoopBackOff (15.14), the loop, ownerReferences and the template hash (15.9), the apply path and Events (15.11), load balancers taking servers in and out (9.23), $do and generators (15.3).

The two objects

apiVersion: apps/v1
kind: Deployment
metadata:
  name: web
spec:
  replicas: 4
  selector:                 # which pods this Deployment manages - IMMUTABLE
    matchLabels:
      app: web
  strategy:
    type: RollingUpdate
    rollingUpdate:
      maxSurge: 1           # at most replicas+1 pods exist during the update
      maxUnavailable: 0     # never fewer than replicas available
  revisionHistoryLimit: 10  # old ReplicaSets kept for rollback
  template:                 # the Pod template; changing it = a new rollout
    metadata:
      labels:
        app: web            # MUST match the selector
    spec:
      containers:
      - name: nginx
        image: nginx:1.27

Field by field:

The Deployment never touches pods directly. It manages ReplicaSets - one per version of the template - and scales them up and down. A ReplicaSet does one thing: keep exactly N pods matching its selector alive (the loop, 15.9). You never create a ReplicaSet yourself: a Deployment gives you everything it does plus updates and history.

The selector must match the template's labels - the API server refuses otherwise:

The Deployment "web" is invalid: spec.template.metadata.labels: Invalid value: map[string]string{"app":"frontend"}: `selector` does not match template `labels`

and the selector cannot be changed after creation (it is immutable):

The Deployment "web" is invalid: spec.selector: Invalid value: v1.LabelSelector{MatchLabels:map[string]string{"app":"web2"}, MatchExpressions:[]v1.LabelSelectorRequirement(nil)}: field is immutable

What triggers a rollout

A rollout is the process of replacing the pods with a new version. Only a change to spec.template starts one - image, environment, labels, annotations inside the template. Changing replicas is a scale, not a rollout: same ReplicaSet, no new version.

k set image deploy/web nginx=nginx:1.28              # container=image
k set env deploy/web LOG_LEVEL=debug
k edit deploy web                                     # change anything in template
k apply -f web.yaml                                   # with a changed template
k rollout restart deploy/web                          # adds a restartedAt annotation

Each version of the template is a revision: revision 1, 2, 3...

RollingUpdate maths

RollingUpdate (the default strategy) replaces pods a few at a time, like taking servers out of a load balancer one by one (9.23). Two knobs:

Both accept a number or a percentage of replicas. maxSurge rounds up, maxUnavailable rounds down. The default is 25% / 25%.

replicas  maxSurge  maxUnavailable  max pods  min available  steps
--------  --------  --------------  --------  -------------  --------------------------
4         25% -> 1  25% -> 1        5         3              fast, loses 1 of capacity
4         1         0               5         4              the Notion question: never
                                                             below 4, needs room for 1
4         0         1               4         3              no extra capacity needed,
                                                             runs at 3 during the update
10        25% -> 3  25% -> 2        13        8
3         25% -> 1  25% -> 0        4         3              (0.75 rounds down to 0)

Read a row: 4 replicas, 25% surge (1, rounded up) and 25% unavailable (1, rounded down) means at most 5 pods at once and at least 3 ready - fast, but you lose a quarter of capacity for a while.

Setting both to 0 is invalid - there would be no way to make progress:

spec.strategy.rollingUpdate.maxUnavailable: Invalid value: intstr.IntOrString{Type:0, IntVal:0, StrVal:""}: may not be 0 when `maxSurge` is 0

With replicas: 4, maxSurge: 1, maxUnavailable: 0 the dance is:

new RS 0->1                     5 pods (4 old + 1 new), wait for the new one to be Ready
old RS 4->3                     4 pods
new RS 1->2                     5 pods ...
... until new=4, old=0

You can watch it in the Deployment's events (k describe deploy web):

Normal  ScalingReplicaSet  40s  deployment-controller  Scaled up replica set web-5c8d7f9b6d from 0 to 1
Normal  ScalingReplicaSet  36s  deployment-controller  Scaled down replica set web-6b8d9c7f5d from 4 to 3
Normal  ScalingReplicaSet  36s  deployment-controller  Scaled up replica set web-5c8d7f9b6d from 1 to 2
...

Two ReplicaSets, two hashes: web-5c8d7f9b6d is the new template, web-6b8d9c7f5d the old one.

maxUnavailable: 0 guarantees no capacity loss, at the cost of needing headroom - the cluster must have room for the surge pods. On a full cluster that surge pod stays Pending and the rollout never moves.

Recreate

strategy: {type: Recreate} scales the old ReplicaSet to 0, waits for its pods to be gone, then scales the new one up. Downtime, guaranteed. Use it only when two versions must never run at once (a database change the old code cannot survive, a program that must be the only copy running).

Watching a rollout

A fresh Deployment to play with (the delete ... --ignore-not-found removes an old one if it exists, so the revision numbers below match), then a change to its template:

$ k delete deploy web --ignore-not-found; k create deployment web --image=nginx:1.27 --replicas=4
deployment.apps/web created
$ k set image deploy/web nginx=nginx:1.28
deployment.apps/web image updated

kubectl rollout status follows a rollout until it finishes:

$ k rollout status deploy/web
Waiting for deployment "web" rollout to finish: 1 out of 4 new replicas have been updated...
Waiting for deployment "web" rollout to finish: 2 out of 4 new replicas have been updated...
Waiting for deployment "web" rollout to finish: 3 out of 4 new replicas have been updated...
Waiting for deployment "web" rollout to finish: 1 old replicas are pending termination...
deployment "web" successfully rolled out

It exits 0 on success and non-zero on failure (1.7) - put it in every pipeline after apply, with --timeout so it cannot hang for ever:

$ k set image deploy/web nginx=nginx:1.99       # a tag that does not exist
deployment.apps/web image updated
$ k rollout status deploy/web --timeout=60s
Waiting for deployment "web" rollout to finish: 2 out of 4 new replicas have been updated...
error: timed out waiting for the condition

k get deploy web summarises the same thing in columns:

columnmeans
READY 3/4ready pods / wanted pods
UP-TO-DATEpods already on the newest template
AVAILABLEpods ready long enough to count as serving
AGEsince the Deployment was created

If the rollout makes no progress for progressDeadlineSeconds (default 600 seconds), the Deployment's Progressing condition flips to False / ProgressDeadlineExceeded and rollout status exits with error: deployment "web" exceeded its progress deadline. Kubernetes does not roll back on its own - that decision is yours, or your pipeline's.

History and rollback

Each ReplicaSet carries a deployment.kubernetes.io/revision annotation (a note, 15.26). kubectl rollout history lists them:

$ k rollout history deploy/web
deployment.apps/web
REVISION  CHANGE-CAUSE
1         <none>
2         <none>
3         <none>

CHANGE-CAUSE - the "why" of each revision - comes from the kubernetes.io/change-cause annotation on the Deployment. The old --record flag used to set it and is deprecated; set it yourself with kubectl annotate. The deployment controller copies it onto the current ReplicaSet, so annotate right after the change (annotate before and the old revision is stamped with the new reason too), or put the annotation in the manifest you apply:

k annotate deploy/web kubernetes.io/change-cause="bump nginx to 1.28 (CHG-4411)"

Look at one revision, and roll back:

$ k rollout history deploy/web --revision=2      # the template of revision 2
k rollout undo deploy/web                        # back to the previous revision
$ k rollout undo deploy/web --to-revision=1      # to a specific one
deployment.apps/web rolled back

A detail that surprises people: undo does not go back in history, it goes forward. Rolling back to revision 1 copies revision 1's template, which creates revision 4 (and revision 1 disappears from the list, because it is 4 now):

$ k rollout history deploy/web
REVISION  CHANGE-CAUSE
2         <none>
3         <none>
4         <none>

k get rs shows the history as ReplicaSets - the old ones sit at 0 replicas:

$ k get rs
NAME             DESIRED   CURRENT   READY   AGE
web-2flq5t9ff5   0         0         0       40s
web-v9vc988nvd   0         0         0       30s
web-vf2s9shm74   4         4         4       48s

DESIRED is the ReplicaSet's replicas, CURRENT how many pods it has, READY how many are ready. Rollback just scales an old one back up.

pause / resume

k rollout pause deploy/web stops the controller acting on template changes. Make several edits (set image, set env...), then k rollout resume deploy/web rolls them out as one revision instead of several.

What you can now do:

Why it helps

Deployments are what you'll ship every application with, and rollouts are where outages happen. Situations: a release with maxUnavailable: 0, maxSurge: 1 never progresses because the cluster has no room for the surge pod; a rollout hits ProgressDeadlineExceeded and nobody notices because the pipeline never ran rollout status; someone expects Kubernetes to roll back automatically (it doesn't). The Notion question "how do you guarantee no capacity loss during an update?" is exactly maxUnavailable: 0. And in the checkout incident you studied, a working undo instead of a stale runbook command would have saved 13 minutes. exam tasks include set image, rollout history and undo.

FAQ

What exactly triggers a new rollout?

Only a change to spec.template: image, env, resources, probes, or labels and annotations inside the template. Changing replicas is just a scale of the current ReplicaSet, with no new revision. kubectl rollout restart works by adding a restartedAt annotation to the template, which counts as a change.

How are maxSurge and maxUnavailable rounded?

Both accept numbers or percentages of replicas; maxSurge rounds up and maxUnavailable rounds down. The default is 25% each: with 4 replicas that's 1 and 1; with 3 replicas it's 1 surge and 0 unavailable. Both can't be 0, since the rollout could never make progress.

Does Kubernetes roll back a failed rollout automatically?

No. If there's no progress for progressDeadlineSeconds (default 600), the Deployment's Progressing condition becomes False with ProgressDeadlineExceeded, and rollout status exits non-zero. The decision to undo is yours or your tooling's, like a CI step or a progressive-delivery tool. Put kubectl rollout status --timeout=... after every apply in pipelines.

Why did rollout undo to revision 1 create revision 4?

Undo doesn't go back in history; it goes forward by copying the old revision's template, creating a new revision. Revision 1's ReplicaSet becomes the current one again with a new revision number, so revision 1 disappears from the history list. The ReplicaSets themselves, visible with kubectl get rs, are the history.

When should I use the Recreate strategy?

Only when two versions must never run at the same time: a schema change the old code can't survive, or a singleton holding an exclusive lock. Recreate scales the old ReplicaSet to 0, waits for its pods to be gone, then scales up the new one, so there is downtime by design.

In an interview Junior

What is the difference between a Deployment and a ReplicaSet?

A ReplicaSet does one thing: keep exactly N pods matching its selector alive. A Deployment manages ReplicaSets - one per version of its pod template - and adds updates and history.

You never create ReplicaSets yourself. Day to day: k set image, k rollout status --timeout (exits non-zero on failure), k rollout history, k rollout undo --to-revision=N. A stuck rollout hits progressDeadlineSeconds, but Kubernetes does not roll back on its own.

Also asked: How do you configure a Deployment for zero-downtime updates? · A rollout is stuck. How do you investigate and recover? · What is the difference between the RollingUpdate and Recreate strategies?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.