OnCallReady

Lesson 25.26 · CI/CD Pipelines & Helm · 20 min read

Hooks, waiting and failure: what Helm does when things go wrong

In plain words

Imagine moving a class to a new classroom. Before the move, someone must clear the new room (a "pre" job). Then the desks are carried over. Only when every desk is in place and the lights work do you announce "moved!". After the announcement, someone hangs the welcome poster (a "post" job). If the move stops halfway, you need to know exactly where: nothing moved yet, or half the desks are already in the new room.

A helm upgrade is that move. It locks the release as pending-upgrade, runs pre-upgrade hooks like a migration Job, applies manifests, waits with --wait until pods are ready, runs post-upgrade hooks, then marks the revision deployed or failed. --rollback-on-failure moves the desks back automatically.

The problem

A new app version often needs something done around the deploy: change the database tables before the new code starts, or check the app works after. And when a deploy fails halfway, you need to know exactly what state the cluster and the release are in - is the old version still serving? Can the next deploy run? This lesson walks through an upgrade step by step.

What you need to know already: 25.19 (release, revision, statuses), 15.24 (Jobs), 15.16 (rolling updates, readiness), 2.3 (SIGTERM vs SIGKILL).

Words for this lesson

The order of an upgrade

helm upgrade is not one API call. In order:

1. render the chart with the merged values          (template/schema errors stop here)
2. write revision N+1 as pending-upgrade            (the lock)
3. run pre-upgrade hooks, by weight, and WAIT for each
4. apply the manifests: create new, patch changed, delete what N had and N+1 has not
5. with --wait: wait until every resource is ready (or --timeout)
6. run post-upgrade hooks, and wait for them
7. mark N+1 deployed, N superseded                  (or N+1 failed)

Step 4 is also where Helm deletes objects the old revision had and the new one no longer renders - the part kubectl apply cannot do. Where it stops decides what state you are in. A failure in 1 leaves nothing behind. A failure in 3 leaves revision N+1 failed and the cluster untouched - the old version keeps serving. A failure in 5 or 6 leaves N+1 failed with the new objects already applied: the Deployment is mid-rollout or stuck, and Kubernetes' own rolling update is what keeps the old pods serving. And if the helm process dies anywhere between 2 and 7, N+1 stays pending-upgrade forever.

Hooks

A hook is an ordinary manifest (usually a Job, 15.24: a pod that runs to completion) with annotations. helm.sh/hook says when it runs, hook-weight orders several hooks, hook-delete-policy says when Helm deletes it:

metadata:
  name: payments-api-migrate
  annotations:
    "helm.sh/hook": pre-install,pre-upgrade
    "helm.sh/hook-weight": "-5"                  # lower runs first; a STRING
    "helm.sh/hook-delete-policy": before-hook-creation,hook-succeeded

Events: pre-install, post-install, pre-upgrade, post-upgrade, pre-rollback, post-rollback, pre-delete, post-delete, test. For Jobs and Pods, Helm waits until they complete; a failed Job (one that hit its backoff limit, the number of retries it is allowed) fails the operation:

Error: UPGRADE FAILED: pre-upgrade hooks failed: resource Job/pay-hooks/payments-api-migrate not ready. status: Failed, message: Job has reached the specified backoff limit

Delete policies:

before-hook-creation   delete the previous run's object before creating it again (the default)
hook-succeeded         delete it once it succeeded
hook-failed            delete it if it failed

before-hook-creation,hook-succeeded is the one you want for migrations: a successful run leaves nothing, a failed run stays so you can read its logs, and the next attempt is not blocked by jobs.batch "x" already exists.

Hook objects are not part of the release: helm uninstall does not delete them, helm get manifest does not show them (helm get hooks does), and they are not rolled back. A migration hook that changes the schema is a one-way door; helm rollback brings back the old code against the new schema. Migrations must be backwards compatible - expand, deploy, contract: first add the new columns while keeping the old ones, then deploy the code that uses them, and only later remove the old ones. Helm cannot help you there.

For a migration, pre-upgrade is the right event: if it fails, nothing new is deployed. A post-upgrade hook is for things that need the new version running (cache warm-up, a smoke check, a notification).

Waiting

Without a wait flag, Helm 4 waits for hooks only and returns as soon as the API server accepted the objects. A Deployment whose new pods crash-loop is reported STATUS: deployed. In a pipeline that is a green run and a broken service.

--wait                    wait until every resource is ready (Deployments rolled out,
                          Services with endpoints, PVCs bound...)
--wait-for-jobs           also wait for Jobs in the chart to complete
--timeout 5m0s            the budget for EACH wait (default 5m)

The error names the object that was not ready and why:

Error: UPGRADE FAILED: resource Deployment/pay-dev/payments-api not ready. status: InProgress, message: Available: 1/2
context deadline exceeded

context deadline exceeded is Go's way of saying "the time ran out". Your --timeout must be longer than the rollout can take: replicas x (image pull + startup + readiness) / maxSurge (maxSurge = how many extra pods a rolling update may start at once, 15.16). A Spring Boot service with a 60s startup, 6 replicas and maxSurge: 1 needs more than 6 minutes; the default 5 minutes makes every release "fail" while actually succeeding a minute later.

Failing safely

--rollback-on-failure     on failure, roll back to the last deployed revision (implies --wait)
--cleanup-on-fail         delete objects this upgrade CREATED if it fails
--atomic                  Helm 4: deprecated alias of --rollback-on-failure

With --rollback-on-failure a failed upgrade ends with:

Error: UPGRADE FAILED: release payments-api failed, and has been rolled back due to rollback-on-failure being set: resource Deployment/pay-dev/payments-api not ready. status: InProgress, message: ...

and helm history shows N+1 failed followed by N+2 Rollback to N. It is the right default for pipelines deploying to shared environments. It is the wrong default when you need to debug: the evidence (the crashing pods) is gone before you look. For dev, fail without rolling back and read the pods.

Statuses and the lock

deployed           current and healthy (as far as Helm knows)
superseded         was deployed, replaced by a newer revision
failed             the operation failed; the previous deployed revision is still "deployed"
pending-install    }
pending-upgrade    }  an operation is running - or its process was killed
pending-rollback   }
uninstalling / uninstalled

Only one operation at a time: any pending-* revision blocks upgrades with

Error: UPGRADE FAILED: another operation (install/upgrade/rollback) is in progress

Ctrl+C or SIGTERM (2.3) makes Helm 4 give up cleanly and mark the revision failed (Release payments-api has been cancelled.). SIGKILL - an OOM-killed or force-stopped CI agent, a cancelled pipeline whose agent was torn down - leaves the pending revision behind. The incident after the next mission is exactly that.

A cancelled or killed operation leaves Helm's record, not the cluster, stuck: the pods keep running whatever was last applied.

Getting out, in plain words: first make sure no helm process is really still running (a live upgrade shows the same status). Then helm history names the last deployed revision, and helm rollback NAME REV to it writes a new revision and releases the lock. Not helm uninstall - that deletes the app's objects to fix a bookkeeping problem.

helm test

Pods annotated "helm.sh/hook": test, usually in templates/tests/, run on demand:

# on a release whose chart has a test hook
helm test payments-api -n pay-dev
NAME: payments-api
...
Phase:          Succeeded

They are smoke tests against the installed release (call the Service, check a health endpoint). A pipeline runs helm test right after helm upgrade --wait; --logs prints the test pods' output.

History

--history-max (default 10 on upgrade) keeps the last N revisions. Rolling back to something older than that is impossible; 0 means unlimited and a namespace full of Secrets, each holding a full compressed copy of the chart. Ten is fine.

What you can now do

Why it helps

The page "deploy stuck, pipeline red, prod half-updated" is where this lesson pays. The step at which the upgrade stopped tells you the state: a failed pre-upgrade migration means the old version is untouched; a failed wait means new objects are applied and the rolling update is what keeps old pods serving; a killed CI agent leaves pending-upgrade and every later deploy fails with another operation is in progress.

Without --wait, a crash-looping Deployment is reported as deployed and the pipeline is green. With the default 5 minute timeout, a slow Spring Boot rollout "fails" while actually succeeding a minute later. And knowing that hooks are not rolled back is what stops you from promising a safe helm rollback after a schema migration. Interviewers love "how do you run DB migrations with Helm?"

FAQ

What does --wait actually wait for?

Without any wait flag, Helm 4 waits only for hooks and returns as soon as the API server accepts the objects. With --wait (the watcher strategy), it waits until every resource is ready by kstatus rules: Deployments rolled out with available replicas, PVCs bound, Services with endpoints, and so on, up to --timeout (default 5m). --wait-for-jobs also waits for Jobs in the chart to complete.

What is the difference between --rollback-on-failure and --cleanup-on-fail?

--rollback-on-failure (formerly --atomic) rolls the release back to the last deployed revision if the upgrade fails, and implies waiting. --cleanup-on-fail only deletes objects that this failed upgrade created; it does not restore changed ones. Rollback-on-failure is a good default for shared environments; in dev you may prefer neither, so the failing pods stay around for you to inspect.

How do I fix "another operation (install/upgrade/rollback) is in progress"?

It means the latest revision is in a pending-* state, usually because the Helm process was killed mid-operation. First make sure nothing is really running. Then check helm history. Options: helm rollback to the last deployed revision, which usually clears it, or, as a last resort, delete or relabel the stuck release Secret sh.helm.release.v1.<name>.v<N>. Ctrl+C or SIGTERM lets Helm 4 mark it failed cleanly; SIGKILL does not.

Are hook resources rolled back or deleted with the release?

No. Hook objects are not part of the release manifest: helm get manifest does not show them (use helm get hooks), helm uninstall does not delete them unless a delete policy does, and rollback does not undo what they did. A migration hook that changed the schema stays changed, so helm rollback runs the old code against the new schema. Migrations must be backwards compatible.

Which hook-delete-policy should a migration Job use?

before-hook-creation,hook-succeeded. A successful migration is cleaned up, a failed one stays so you can read its logs with kubectl logs job/..., and the next attempt deletes the old Job first instead of failing with jobs.batch "x" already exists. before-hook-creation alone is the default. hook-failed deletes failures too, which throws away the evidence you need.

In an interview Mid

A helm upgrade in the pipeline timed out. What state is the release in, and what do you do?

It depends where helm upgrade stopped. The order is: render, write revision N+1 as pending-upgrade (the lock), pre-upgrade hooks, apply the manifests, wait, post-upgrade hooks, mark deployed.

helm history tells you which case you are in. For shared environments --rollback-on-failure returns to the last good revision automatically - but removes the evidence, so not in dev.

Also asked: How do you run database migrations as part of a Helm deployment? · What does --rollback-on-failure do, and when would you not use it? · Why is helm rollback not enough to undo a database migration?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.