The problem
Argo CD applies a new version, and the Deployment replaces the pods. For most services that is fine. But sometimes you need more: a rollback that takes one second instead of another rollout, a full test of the new version in production before any customer reaches it, or only 5% of customers on the new version until you trust it. Those are delivery strategies - the ways a new version replaces the old one.
What you need to know already: 17.24 (rolling updates: maxSurge, maxUnavailable, the rollout verbs), 17.20 (readiness probes), 16.1 (Services select pods by label), 16.19 (Ingress and the ingress controller), 0.1 (the four golden signals: latency, traffic, errors, saturation), 26.4 (Argo CD health).
RollingUpdate, precisely
The Deployment default (17.24). Two knobs decide the shape of the rollout:
spec:
replicas: 4
minReadySeconds: 10 # a new pod must stay Ready this long to count as available
progressDeadlineSeconds: 600 # no progress for 10 min -> Progressing=False (ProgressDeadlineExceeded)
strategy:
type: RollingUpdate
rollingUpdate:
maxSurge: 25% # extra pods allowed above replicas (rounded UP)
maxUnavailable: 25% # pods allowed below replicas (rounded DOWN)
With 4 replicas: 25% of 4 = 1, so at most 5 pods exist and at least 3 are available at every moment.
maxSurge: 0, maxUnavailable: 1never needs spare capacity (namespaces tight on quota).maxSurge: 1, maxUnavailable: 0never drops capacity (the safe default for services where slowness hurts).- Both zero is rejected: the rollout could never move.
Without readiness probes none of this means anything - "available" is then just "started".
Recreate kills all old pods, then starts new ones: downtime, but never two versions at once - for apps that cannot run side by side (a single consumer of a queue, an exclusive lock, a database change the old version cannot live with).
A gotcha: switching an existing Deployment to Recreate with kubectl apply (or Argo CD) fails with spec.strategy.rollingUpdate: Forbidden: may not be specified when strategy type is 'Recreate'. The API server filled in rollingUpdate defaults when the Deployment was created; declare rollingUpdate: null to remove it.
What rolling cannot give you: an instant rollback (it is another rollout), a test of the new version with real traffic before users get it, or control over how much traffic the new version gets (it is proportional to pod count, and only once pods are ready).
Blue-green
Run the new version (green) next to the old (blue) at full size. Test green through a separate preview Service. Then switch the production Service's selector (16.1) from blue's label to green's:
kind: Service
metadata: {name: payments-api}
spec:
selector: {app.kubernetes.io/name: payments-api, track: green} # was: blue
- The switch happens at once for new connections.
- Rollback is switching back - blue is still running.
- You can test green in production conditions first.
Costs: double capacity during the release; long-lived connections stay on blue until they close; and both versions share the database, so the schema must work for both.
Canary
Send a small share of real traffic to the new version (the canary, after the bird miners took underground as an early warning), watch it, then increase. The old version is called stable. Two ways:
- By replica ratio: a stable Deployment (3 pods) and a canary Deployment (1 pod) behind one Service = roughly 25% canary. Simple, coarse (1 pod out of N), and the ratio depends on pod counts.
- By traffic split: something in front of the pods sends exactly 5%, 20%, 50%... regardless of replica counts, and can route by header ("only internal testers"). That something is the ingress controller (NGINX canary annotations, 16.19), Gateway API route weights (16.26), or a service mesh (Istio, Linkerd: a proxy container added to every pod that controls traffic between services).
Promotion and rollback are decided by measurements: the error rate and latency of the canary compared with stable - the golden signals from 0.1. Automated canary analysis queries a monitoring system for those numbers at each step and aborts on its own if the canary is worse.
Later (Ch 27): the monitoring system those numbers come from is Prometheus; the analysis step then runs a Prometheus query.
Argo Rollouts
Argo Rollouts is a separate controller (from the same project as Argo CD) that replaces the Deployment with a Rollout object which knows these strategies:
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata: {name: payments-api}
spec:
replicas: 4
strategy:
canary:
canaryService: payments-api-canary # a Service that selects only canary pods
stableService: payments-api # a Service that selects only stable pods
trafficRouting:
nginx: {stableIngress: payments-api} # split traffic at the NGINX ingress
steps:
- setWeight: 10 # 10% of traffic to the canary
- pause: {duration: 5m} # wait 5 minutes
- analysis:
templates: [{templateName: error-rate}] # measure the error rate, fail if > 1%
- setWeight: 50
- pause: {} # wait for a human: kubectl argo rollouts promote
template: { ... same as a Deployment's ... }
blueGreen: works the same way with activeService, previewService and autoPromotionEnabled: false (wait for a human before switching). kubectl argo rollouts get rollout payments-api --watch shows the steps live; kubectl argo rollouts abort payments-api sends everything back to stable. Argo CD shows a Rollout paused at a step as Suspended (26.4). Flagger is the equivalent in the Flux world.
Feature flags are not a deployment strategy
A feature flag is a switch in the application's config: the new code ships turned off ("dark") and is turned on per user or percentage while it runs, rolled back by flipping the switch. It separates deploying code from releasing a feature. It complements the strategies above - canary tests the new build, flags test the feature - and it has its own debt (old flags must be removed).
Choosing
rolling default; stateless services; cheap; rollback = another rollout
recreate cannot run two versions at once; accepts downtime
blue-green need instant rollback and a full test in prod before switching;
can afford double capacity
canary want real-traffic evidence before full exposure; have measurements
(and ideally traffic splitting) to decide promotion
Every one of them runs two versions against the same database and the same clients at some point. The expand / contract discipline is what makes any strategy safe: expand first (add new columns before any code uses them; APIs accept both old and new shapes), contract later (remove only after no running version reads them). Without it, even a perfect canary breaks the stable pods.
What you can now do
- Work out how many pods a rolling update may add or remove at once.
- Run a blue-green switch by changing a Service selector, and say what it costs.
- Explain canary by replica ratio vs by traffic split, and what decides promotion.