OnCallReady

Lesson 26.26 · GitOps, Argo CD & Delivery · 12 min read

Delivery strategies: rolling, blue-green, canary

In plain words

Imagine a school changing the lunch menu. Rolling: swap the food at one serving counter at a time, so there is always food, and for a while some kids get the old menu and some the new. Blue-green: cook the whole new menu in a second kitchen, taste it, then open the doors to the new kitchen at once, keeping the old one ready in case. Canary: give the new menu to one table first, watch whether anyone gets a stomach ache, then give it to more tables.

In Kubernetes, rolling is the Deployment default, shaped by maxSurge and maxUnavailable. Blue-green switches the Service selector from track: blue to track: green. Canary sends a small share of traffic to the new version and decides by metrics; Argo Rollouts automates steps like setWeight: 10, pause, and an analysis query.

The problem

Argo CD applies a new version, and the Deployment replaces the pods. For most services that is fine. But sometimes you need more: a rollback that takes one second instead of another rollout, a full test of the new version in production before any customer reaches it, or only 5% of customers on the new version until you trust it. Those are delivery strategies - the ways a new version replaces the old one.

What you need to know already: 17.24 (rolling updates: maxSurge, maxUnavailable, the rollout verbs), 17.20 (readiness probes), 16.1 (Services select pods by label), 16.19 (Ingress and the ingress controller), 0.1 (the four golden signals: latency, traffic, errors, saturation), 26.4 (Argo CD health).

RollingUpdate, precisely

The Deployment default (17.24). Two knobs decide the shape of the rollout:

spec:
  replicas: 4
  minReadySeconds: 10            # a new pod must stay Ready this long to count as available
  progressDeadlineSeconds: 600   # no progress for 10 min -> Progressing=False (ProgressDeadlineExceeded)
  strategy:
    type: RollingUpdate
    rollingUpdate:
      maxSurge: 25%              # extra pods allowed above replicas   (rounded UP)
      maxUnavailable: 25%        # pods allowed below replicas         (rounded DOWN)

With 4 replicas: 25% of 4 = 1, so at most 5 pods exist and at least 3 are available at every moment.

Without readiness probes none of this means anything - "available" is then just "started".

Recreate kills all old pods, then starts new ones: downtime, but never two versions at once - for apps that cannot run side by side (a single consumer of a queue, an exclusive lock, a database change the old version cannot live with).

A gotcha: switching an existing Deployment to Recreate with kubectl apply (or Argo CD) fails with spec.strategy.rollingUpdate: Forbidden: may not be specified when strategy type is 'Recreate'. The API server filled in rollingUpdate defaults when the Deployment was created; declare rollingUpdate: null to remove it.

What rolling cannot give you: an instant rollback (it is another rollout), a test of the new version with real traffic before users get it, or control over how much traffic the new version gets (it is proportional to pod count, and only once pods are ready).

Blue-green

Run the new version (green) next to the old (blue) at full size. Test green through a separate preview Service. Then switch the production Service's selector (16.1) from blue's label to green's:

kind: Service
metadata: {name: payments-api}
spec:
  selector: {app.kubernetes.io/name: payments-api, track: green}   # was: blue

Costs: double capacity during the release; long-lived connections stay on blue until they close; and both versions share the database, so the schema must work for both.

Canary

Send a small share of real traffic to the new version (the canary, after the bird miners took underground as an early warning), watch it, then increase. The old version is called stable. Two ways:

Promotion and rollback are decided by measurements: the error rate and latency of the canary compared with stable - the golden signals from 0.1. Automated canary analysis queries a monitoring system for those numbers at each step and aborts on its own if the canary is worse.

Later (Ch 27): the monitoring system those numbers come from is Prometheus; the analysis step then runs a Prometheus query.

Argo Rollouts

Argo Rollouts is a separate controller (from the same project as Argo CD) that replaces the Deployment with a Rollout object which knows these strategies:

apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata: {name: payments-api}
spec:
  replicas: 4
  strategy:
    canary:
      canaryService: payments-api-canary        # a Service that selects only canary pods
      stableService: payments-api               # a Service that selects only stable pods
      trafficRouting:
        nginx: {stableIngress: payments-api}    # split traffic at the NGINX ingress
      steps:
        - setWeight: 10                         # 10% of traffic to the canary
        - pause: {duration: 5m}                 # wait 5 minutes
        - analysis:
            templates: [{templateName: error-rate}]    # measure the error rate, fail if > 1%
        - setWeight: 50
        - pause: {}                             # wait for a human: kubectl argo rollouts promote
  template: { ... same as a Deployment's ... }

blueGreen: works the same way with activeService, previewService and autoPromotionEnabled: false (wait for a human before switching). kubectl argo rollouts get rollout payments-api --watch shows the steps live; kubectl argo rollouts abort payments-api sends everything back to stable. Argo CD shows a Rollout paused at a step as Suspended (26.4). Flagger is the equivalent in the Flux world.

Feature flags are not a deployment strategy

A feature flag is a switch in the application's config: the new code ships turned off ("dark") and is turned on per user or percentage while it runs, rolled back by flipping the switch. It separates deploying code from releasing a feature. It complements the strategies above - canary tests the new build, flags test the feature - and it has its own debt (old flags must be removed).

Choosing

rolling          default; stateless services; cheap; rollback = another rollout
recreate         cannot run two versions at once; accepts downtime
blue-green       need instant rollback and a full test in prod before switching;
                 can afford double capacity
canary           want real-traffic evidence before full exposure; have measurements
                 (and ideally traffic splitting) to decide promotion

Every one of them runs two versions against the same database and the same clients at some point. The expand / contract discipline is what makes any strategy safe: expand first (add new columns before any code uses them; APIs accept both old and new shapes), contract later (remove only after no running version reads them). Without it, even a perfect canary breaks the stable pods.

What you can now do

Why it helps

The strategy decides how bad a bad release is. A rolling update with maxUnavailable: 25% and no readiness probe can take down a quarter of capacity instantly; maxSurge: 1, maxUnavailable: 0 never drops capacity. A canary at 10% with automated analysis turns a full outage into a small, automatically aborted blip. Blue-green gives instant rollback when a regulator asks for it, at the cost of double capacity.

You will pick and configure these for services, debug a Recreate switch that fails with may not be specified when strategy type is 'Recreate', and see Argo CD show a paused Rollout as Suspended. And every strategy runs two versions against one database at some point, which is why expand and contract comes up in every senior interview about deployments.

FAQ

What do maxSurge and maxUnavailable mean?

maxSurge is how many extra pods may exist above replicas during a rollout (percentages round up). maxUnavailable is how many may be missing below replicas (rounding down). With 4 replicas and 25% each, at most 5 pods exist and at least 3 are available. maxSurge: 1, maxUnavailable: 0 never drops capacity; maxSurge: 0, maxUnavailable: 1 needs no spare capacity. Both zero is rejected.

What is the difference between blue-green and canary?

Blue-green runs the new version at full size next to the old, tests it through a preview Service, then switches all traffic at once by changing the Service selector; rollback is switching back. Canary sends a small share of real traffic to the new version first and increases it step by step based on metrics. Blue-green gives instant cutover and rollback; canary limits the blast radius and gives real-traffic evidence.

Can I do a canary with plain Kubernetes?

Roughly. Run a stable Deployment and a canary Deployment with the same labels behind one Service: with 3 stable pods and 1 canary, about 25% of requests hit the canary. It is coarse, the percentage depends on pod counts, and you must watch metrics and promote by hand. For precise weights like 5%, or routing by header, you need an ingress controller, Gateway API or a service mesh, typically driven by Argo Rollouts or Flagger.

When should I use Recreate?

When two versions cannot run at the same time: a singleton consumer that must not process messages twice, an exclusive lock, or a schema change the old version cannot tolerate. Recreate stops all old pods before starting new ones, so there is downtime. When switching an existing Deployment to Recreate, also set rollingUpdate: null, because the API server defaulted that block and it is forbidden with type: Recreate.

Are feature flags a deployment strategy?

No, they complement one. A deployment strategy controls how a new binary replaces the old one. A feature flag controls whether a feature in already-deployed code is active, per user, group or percentage, and can be switched off without a redeploy. Canary tests the binary, flags test the feature. Flags carry their own debt: stale flags must be removed or the code fills with dead branches.

In an interview Mid

Explain rolling update, blue-green and canary deployments.

All of them run two versions against one database at some point, so schema changes must be expand / contract. Feature flags are separate: they release a feature, not a build.

Also asked: How would you configure a rolling update for a latency-sensitive service with a slow startup? · What decides whether a canary is promoted or rolled back? · Why does every deployment strategy need backwards-compatible database changes?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.