The problem
A new app version often needs something done around the deploy: change the database tables before the new code starts, or check the app works after. And when a deploy fails halfway, you need to know exactly what state the cluster and the release are in - is the old version still serving? Can the next deploy run? This lesson walks through an upgrade step by step.
What you need to know already: 25.19 (release, revision, statuses), 15.24 (Jobs), 15.16 (rolling updates, readiness), 2.3 (SIGTERM vs SIGKILL).
Words for this lesson
- hook - a Kubernetes object in the chart that Helm runs at a set moment (before an upgrade, after an install...), instead of applying it with the rest.
- migration - a script that changes a database's structure (tables, columns) to what the new code expects. Flyway and Liquibase are the common Java tools for it.
- lock (Helm sense) - the
pending-*status that stops a second operation on the same release. - smoke test - a quick check that the installed app basically works.
The order of an upgrade
helm upgrade is not one API call. In order:
1. render the chart with the merged values (template/schema errors stop here)
2. write revision N+1 as pending-upgrade (the lock)
3. run pre-upgrade hooks, by weight, and WAIT for each
4. apply the manifests: create new, patch changed, delete what N had and N+1 has not
5. with --wait: wait until every resource is ready (or --timeout)
6. run post-upgrade hooks, and wait for them
7. mark N+1 deployed, N superseded (or N+1 failed)
Step 4 is also where Helm deletes objects the old revision had and the new one no longer renders - the part kubectl apply cannot do. Where it stops decides what state you are in. A failure in 1 leaves nothing behind. A failure in 3 leaves revision N+1 failed and the cluster untouched - the old version keeps serving. A failure in 5 or 6 leaves N+1 failed with the new objects already applied: the Deployment is mid-rollout or stuck, and Kubernetes' own rolling update is what keeps the old pods serving. And if the helm process dies anywhere between 2 and 7, N+1 stays pending-upgrade forever.
Hooks
A hook is an ordinary manifest (usually a Job, 15.24: a pod that runs to completion) with annotations. helm.sh/hook says when it runs, hook-weight orders several hooks, hook-delete-policy says when Helm deletes it:
metadata:
name: payments-api-migrate
annotations:
"helm.sh/hook": pre-install,pre-upgrade
"helm.sh/hook-weight": "-5" # lower runs first; a STRING
"helm.sh/hook-delete-policy": before-hook-creation,hook-succeeded
Events: pre-install, post-install, pre-upgrade, post-upgrade, pre-rollback, post-rollback, pre-delete, post-delete, test. For Jobs and Pods, Helm waits until they complete; a failed Job (one that hit its backoff limit, the number of retries it is allowed) fails the operation:
Error: UPGRADE FAILED: pre-upgrade hooks failed: resource Job/pay-hooks/payments-api-migrate not ready. status: Failed, message: Job has reached the specified backoff limit
Delete policies:
before-hook-creation delete the previous run's object before creating it again (the default)
hook-succeeded delete it once it succeeded
hook-failed delete it if it failed
before-hook-creation,hook-succeeded is the one you want for migrations: a successful run leaves nothing, a failed run stays so you can read its logs, and the next attempt is not blocked by jobs.batch "x" already exists.
Hook objects are not part of the release: helm uninstall does not delete them, helm get manifest does not show them (helm get hooks does), and they are not rolled back. A migration hook that changes the schema is a one-way door; helm rollback brings back the old code against the new schema. Migrations must be backwards compatible - expand, deploy, contract: first add the new columns while keeping the old ones, then deploy the code that uses them, and only later remove the old ones. Helm cannot help you there.
For a migration, pre-upgrade is the right event: if it fails, nothing new is deployed. A post-upgrade hook is for things that need the new version running (cache warm-up, a smoke check, a notification).
Waiting
Without a wait flag, Helm 4 waits for hooks only and returns as soon as the API server accepted the objects. A Deployment whose new pods crash-loop is reported STATUS: deployed. In a pipeline that is a green run and a broken service.
--wait wait until every resource is ready (Deployments rolled out,
Services with endpoints, PVCs bound...)
--wait-for-jobs also wait for Jobs in the chart to complete
--timeout 5m0s the budget for EACH wait (default 5m)
The error names the object that was not ready and why:
Error: UPGRADE FAILED: resource Deployment/pay-dev/payments-api not ready. status: InProgress, message: Available: 1/2
context deadline exceeded
context deadline exceeded is Go's way of saying "the time ran out". Your --timeout must be longer than the rollout can take: replicas x (image pull + startup + readiness) / maxSurge (maxSurge = how many extra pods a rolling update may start at once, 15.16). A Spring Boot service with a 60s startup, 6 replicas and maxSurge: 1 needs more than 6 minutes; the default 5 minutes makes every release "fail" while actually succeeding a minute later.
Failing safely
--rollback-on-failure on failure, roll back to the last deployed revision (implies --wait)
--cleanup-on-fail delete objects this upgrade CREATED if it fails
--atomic Helm 4: deprecated alias of --rollback-on-failure
With --rollback-on-failure a failed upgrade ends with:
Error: UPGRADE FAILED: release payments-api failed, and has been rolled back due to rollback-on-failure being set: resource Deployment/pay-dev/payments-api not ready. status: InProgress, message: ...
and helm history shows N+1 failed followed by N+2 Rollback to N. It is the right default for pipelines deploying to shared environments. It is the wrong default when you need to debug: the evidence (the crashing pods) is gone before you look. For dev, fail without rolling back and read the pods.
Statuses and the lock
deployed current and healthy (as far as Helm knows)
superseded was deployed, replaced by a newer revision
failed the operation failed; the previous deployed revision is still "deployed"
pending-install }
pending-upgrade } an operation is running - or its process was killed
pending-rollback }
uninstalling / uninstalled
Only one operation at a time: any pending-* revision blocks upgrades with
Error: UPGRADE FAILED: another operation (install/upgrade/rollback) is in progress
Ctrl+C or SIGTERM (2.3) makes Helm 4 give up cleanly and mark the revision failed (Release payments-api has been cancelled.). SIGKILL - an OOM-killed or force-stopped CI agent, a cancelled pipeline whose agent was torn down - leaves the pending revision behind. The incident after the next mission is exactly that.
A cancelled or killed operation leaves Helm's record, not the cluster, stuck: the pods keep running whatever was last applied.
Getting out, in plain words: first make sure no helm process is really still running (a live upgrade shows the same status). Then helm history names the last deployed revision, and helm rollback NAME REV to it writes a new revision and releases the lock. Not helm uninstall - that deletes the app's objects to fix a bookkeeping problem.
helm test
Pods annotated "helm.sh/hook": test, usually in templates/tests/, run on demand:
# on a release whose chart has a test hook
helm test payments-api -n pay-dev
NAME: payments-api
...
Phase: Succeeded
They are smoke tests against the installed release (call the Service, check a health endpoint). A pipeline runs helm test right after helm upgrade --wait; --logs prints the test pods' output.
History
--history-max (default 10 on upgrade) keeps the last N revisions. Rolling back to something older than that is impossible; 0 means unlimited and a namespace full of Secrets, each holding a full compressed copy of the chart. Ten is fine.
What you can now do
- List the seven steps of
helm upgradeand say what state a failure at each one leaves. - Write a migration hook with the right event and delete policy.
- Choose
--wait,--timeoutand--rollback-on-failure, and explain thepending-upgradelock.