OnCallReady

Lesson 26.9 · GitOps, Argo CD & Delivery · 11 min read

Sync policies: automated, prune, self-heal, waves and hooks

In plain words

Imagine a robot tidying your room to match a photo of how it should look. Three switches on the robot. "Automatic": it tidies whenever the photo changes, otherwise it waits for you to press go. "Throw away": if something is in the room but not in the photo, it bins it; without this switch it only points at it. "Undo meddling": if your sister moves things around, it puts them back even though the photo did not change.

Those are Argo CD's automated, prune and selfHeal. On top, the robot works in order: first the PreSync jobs (like a migration), then the main objects in waves (CRDs and Namespaces first, the app later, each wave waiting until healthy), then PostSync jobs such as a smoke test.

The problem

So far you pressed Sync yourself (26.5, 26.6). That still leaves a human in the loop for every change, and drift is only reported, never fixed. You want Argo CD to apply each commit by itself, delete what was removed from git, and undo changes made behind its back - and you want control over the order things are applied in (a database migration before the new app version).

What you need to know already: 26.4 (Application, sync, sync and health status), 26.6 (drift and self-heal seen once), 25.26 (Helm hooks: a Job that runs before or after an upgrade), 15.24 (Jobs), 15.26 (annotations).

Three switches

syncPolicy:
  automated:
    prune: true        # delete live objects that are no longer in git
    selfHeal: true     # re-sync when the LIVE state drifts, not only when git changes
    allowEmpty: false  # refuse to sync if git renders nothing (protects against a bad path)

Auto-sync does not retry a commit that already failed: if the sync of commit a1b2c3 fails, it stays failed until a new commit (or a manual sync). Add retry: if short-lived failures (a component not ready yet, a CRD still installing) should be retried:

  retry:
    limit: 5                                             # at most 5 more tries
    backoff: { duration: 5s, factor: 2, maxDuration: 3m }  # wait 5s, 10s, 20s... up to 3m

The same switches from the shell: argocd app set APP --sync-policy automated --auto-prune --self-heal.

Sync options

Sync options tune how a sync applies things. They go in syncPolicy.syncOptions:

CreateNamespace=true          create destination.namespace if missing
PruneLast=true                prune after everything else is applied and healthy
ApplyOutOfSyncOnly=true       only apply what differs (big apps: faster, less load on the API)
ServerSideApply=true          use kubectl apply --server-side (the API server tracks which
                              tool owns which field; avoids the 262 KiB size limit of the
                              last-applied annotation - big CRDs, big ConfigMaps)
Replace=true                  kubectl replace/create instead of apply (last resort)
RespectIgnoreDifferences=true do not overwrite fields listed in ignoreDifferences on sync (26.12)
Validate=false                skip schema validation

Per object, as annotations:

Phases and hooks

A sync runs in phases, in order: PreSync -> Sync -> PostSync, and SyncFail only if it failed. Your normal manifests are in the Sync phase.

A sync hook is any manifest - usually a Job - annotated to run in another phase:

metadata:
  annotations:
    argocd.argoproj.io/hook: PreSync              # PreSync | Sync | PostSync | SyncFail | Skip
    argocd.argoproj.io/hook-delete-policy: BeforeHookCreation   # | HookSucceeded | HookFailed

The delete policy says when Argo CD deletes the old hook Job: BeforeHookCreation = just before creating it again on the next sync (so you can still read the last run's logs); HookSucceeded / HookFailed = right after it succeeds / fails.

A failed PreSync Job stops the sync before any Deployment changes - the database migration pattern from 25.26. When Argo CD renders a Helm chart, it maps Helm's pre-install/pre-upgrade hooks to PreSync and post-* to PostSync. helm.sh/hook: test is ignored.

Waves

Within a phase, sync waves order the apply:

    argocd.argoproj.io/sync-wave: "-1"    # lower first; default 0; must be a STRING (quoted)

Argo CD applies wave -1, waits until those objects are healthy, then wave 0, and so on. Typical use: CRDs and Namespaces at -2, ConfigMaps/Secrets at -1, the app at 0, a smoke-test Job (a quick "does it answer?" check) as a PostSync hook. Inside one wave, objects go in kind order (Namespaces, then ConfigMaps/Secrets, then Services, then Deployments...).

Deleting an app

A finalizer is a marker on a Kubernetes object that says "someone must clean up before this object may disappear" (the API waits until it is removed). An Application with the finalizer resources-finalizer.argocd.argoproj.io deletes everything it deployed when you delete it (cascade). Without it, deleting the Application leaves the objects running, orphaned. argocd app delete cascades by default; kubectl delete application only if the finalizer is set. Know which one your team uses before you "just recreate the app".

What you can now do

Why it helps

Sync policy decides what happens during an incident. With selfHeal on, a hotfix with kubectl edit disappears within seconds, which is exactly the incident in this chapter. Without prune, deleting a file from git leaves a zombie Service running and the app OutOfSync forever. Auto-sync does not retry a failed commit unless you set retry:, so a transient webhook failure can leave dev stuck until the next push.

You will choose these per environment on a platform team: auto-sync plus prune plus selfHeal for dev, often auto-sync without auto-prune for prod, Prune=false on PVCs. Waves and hooks order CRDs before custom resources and migrations before the app. And knowing whether your Applications carry the resources finalizer tells you whether "just delete and recreate the app" also deletes production workloads.

FAQ

What is the difference between prune and selfHeal?

Prune deals with objects that were removed from git: with it, Argo CD deletes them from the cluster; without it, they keep running and the app shows OutOfSync with "requires pruning". SelfHeal deals with live changes: when someone edits an object in the cluster, Argo CD syncs it back to git even though git did not change. Without selfHeal, a manual edit stays until the next commit is synced.

Why doesn't auto-sync retry my failed sync?

Automated sync attempts each commit once. If the sync of commit a1b2c3 fails, Argo CD will not try the same revision again automatically, to avoid hammering the cluster; it waits for a new commit or a manual sync. For transient failures, like a CRD still installing or an admission webhook not ready, add retry: with a limit and backoff to the sync policy.

What are sync waves and how do they differ from hooks?

Waves order normal resources within a phase: the annotation argocd.argoproj.io/sync-wave: "-1" applies lower waves first, and Argo CD waits until each wave is healthy before the next. Hooks are resources that run in a specific phase, PreSync, Sync, PostSync or SyncFail, usually Jobs, with their own delete policies. A migration is a PreSync hook; ordering CRDs before the app is a wave.

Do Helm hooks work when Argo CD renders a chart?

Partly. Argo CD renders charts with helm template and applies them itself, so there is no Helm release. It maps Helm's pre-install and pre-upgrade hooks to PreSync, post-install and post-upgrade to PostSync, and ignores helm.sh/hook: test. helm list shows nothing for Argo CD-managed charts, and helm rollback does not apply; history lives in Argo CD and git.

What happens when I delete an Application?

It depends on the finalizer. With resources-finalizer.argocd.argoproj.io, deleting the Application also deletes every resource it deployed (cascade). Without it, the objects keep running, orphaned. argocd app delete cascades by default; kubectl delete application only if the finalizer is present. Check before recreating an app in prod, and mark critical objects like PVCs with Delete=false.

In an interview Mid

What do automated sync, prune and self-heal do in Argo CD?

syncPolicy:
  automated:
    prune: true
    selfHeal: true

Two details: auto-sync does not retry a commit that already failed (add retry: with backoff for transient failures), and allowEmpty: false refuses to sync a path that suddenly renders nothing.

Ordering is separate: PreSync hooks (a migration Job) run before the apply, and sync waves apply lower waves first and wait for them to be healthy.

Also asked: How do you order resources during an Argo CD sync, for example a migration before the app? · What sync policies would you choose for dev and for prod, and why? · What happens to the deployed objects when you delete an Argo CD Application?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.