OnCallReady

Lesson 26.1 · GitOps, Argo CD & Delivery · 11 min read

GitOps: git as the only way in

In plain words

Think of a Lego model built from the instruction booklet. A helper keeps comparing the model with the booklet, all day long. If your little brother pulls off a piece or adds a random one, the helper notices and puts it back to what the booklet shows. If you want the model to change, you do not touch the model; you change the booklet, and write in it who changed what and why.

GitOps works the same way. The booklet is the config repo (payments-gitops), with one directory per environment. The helper is Argo CD inside the cluster, pulling git and reconciling. The pipeline's only job is to write a new line in the booklet: the new image digest. Promotion is a pull request; rollback is git revert of that commit.

The problem

In chapter 25 the pipeline ran helm upgrade against the cluster. That works, but three things hurt. Every pipeline (and every agent that runs one) holds credentials that can change prod. The cluster is simply whatever the last run made it. And anyone with kubectl can quietly change prod, and nobody notices until the next deploy overwrites it - or does not. GitOps is the fix for all three.

What you need to know already: 25.1 (delivery, CI and CD), 25.2 (git: commits, branches, pull requests, revert), 25.11 (promotion by digest), 25.19 (Helm), 15.9 (the reconciliation loop), 15.40 (declarative YAML and kubectl apply).

Push and pull

What chapter 25 did is push-based CD: something outside the cluster (the pipeline) pushes changes in. It needs cluster credentials, and the cluster's state is the sum of whatever runs happened.

GitOps turns it around: git holds the desired state, and a program inside the cluster pulls it and applies it. The desired state is what you want running (the YAML in git); the live state is what actually runs in the cluster.

The four principles

The OpenGitOps project (a group that wrote down what "GitOps" means) defines four principles:

  1. Declarative. The desired state is written as data (YAML, a Helm chart and values, a kustomize overlay - 26.2), not as a list of commands to run.
  2. Versioned and immutable. That data is stored somewhere that keeps every version and never rewrites old ones: git. Every change has an author, a review, a timestamp and a revert.
  3. Pulled automatically. Software agents in (or next to) the cluster pull the desired state. Nothing outside pushes into the cluster.
  4. Continuously reconciled. The agents keep comparing live state with desired state and fix the difference - not only at deploy time, always. This is the reconciliation loop from 15.9, with git as the source.

When live and desired differ, that difference is called drift. Fixing it is reconciling.

The agent is Argo CD (lesson 26.4) or a similar tool called Flux. The pipeline's job shrinks: build, test, push an image, and commit the new image digest to the config repo. Argo CD notices the commit and applies it.

Why it is worth it

                       push (ch25)                       pull (GitOps)
who has cluster creds  every pipeline, every agent       only the in-cluster controller
source of truth        the cluster + the last run        git
drift                  silent until the next deploy      detected in minutes, optionally reverted
audit                  pipeline logs (kept N days)       git log / pull request history
rollback               re-run an old pipeline            git revert (or argocd rollback)
disaster recovery      re-run every pipeline, in order   point a new cluster at the repo

The credentials line is the one security teams care about. In a bank, "the CI system can kubectl apply anything to prod" is an audit finding; "a controller in the cluster applies what was merged to a protected branch, after review" is a control (a check the auditors accept).

What it costs: another moving part to run and understand, a second repo, and a new way to fail - the controller and git can disagree for reasons that have nothing to do with your change. You meet two of them as incidents in this chapter.

Two repos

pay/payments-api       the APP repo:    code, Dockerfile, tests, the pipeline
                       -> produces images (and maybe the chart)
pay/payments-gitops    the CONFIG repo: what runs where - base + overlays / values per env
                       -> read by Argo CD

Why split them:

Environments are directories, not branches

base/                    the service, the same everywhere
overlays/dev/            what differs in dev
overlays/prod/           what differs in prod

Branch-per-environment (dev, staging, prod branches) looks natural and fails in practice: promoting becomes a merge that drags unrelated changes along, branches drift apart, and "what is different between staging and prod" becomes a merge puzzle instead of diff -r overlays/staging overlays/prod (diff -r compares two directories file by file). Every Argo CD Application points at the main branch, each at a different path.

Promotion and rollback

Promotion to prod (25.11: moving the same built image to the next environment) is a pull request that changes one line:

 images:
   - name: registry.lab/pay/payments-api
-    digest: sha256:4be1...
+    digest: sha256:9c3e...

The approval is the pull request review (plus branch protection), the audit record is the merge commit, the deploy is Argo CD noticing it. Rollback is git revert of that commit (25.2) - a new commit, reviewed like any other. Tools such as Argo CD Image Updater, Flux image automation or Renovate can open these pull requests for you; for dev many teams let the pipeline commit directly.

What GitOps does not solve

What you can now do

Why it helps

This lesson is the "why" you will defend in design reviews and interviews. The credentials argument is the one security teams care about: in push-based CD every pipeline agent can kubectl apply to prod; with GitOps only the in-cluster controller can. The drift argument is the one on-call cares about: a hand edit made at 2am is visible within minutes instead of surfacing at the next deploy.

It also sets the rules you will enforce on a platform team: two repos (app and config) so a digest commit does not retrigger the build, directories not branches for environments, and CODEOWNERS on overlays/prod/. When someone asks "how do we roll back?", your answer is a reviewed git revert, not re-running an old pipeline whose logs may already have expired.

FAQ

What are the four OpenGitOps principles?

Declarative: desired state is data, not a script of commands. Versioned and immutable: it is stored in a system that keeps every version, in practice git. Pulled automatically: agents pull the desired state rather than something outside pushing into the cluster. Continuously reconciled: agents keep comparing actual with desired state and correct differences at all times, not only at deploy time.

Why not let the pipeline commit the new digest to the app repo?

Because that commit would trigger the app pipeline again, a loop you have to guard against with skip-ci tricks. It also mixes deployment history with code history and gives the same people rights over both. A separate config repo means a digest commit triggers only Argo CD, git log overlays/prod is a clean deployment list, and prod config can have its own reviewers.

How is rollback done in GitOps?

With git revert of the promotion commit: a new commit restoring the previous digest, reviewed like any change, which Argo CD then syncs. The history shows both the deploy and its reversal. argocd app rollback exists but only works with auto-sync off, and it leaves git and cluster disagreeing, so the next sync would redeploy the bad version unless git is fixed too.

Does GitOps handle database migrations?

Not by itself. Argo CD can run a migration Job as a PreSync hook, so a failed migration stops the sync, but ordering and compatibility are still your problem. A rollback via git revert restores the old image, not the old schema. Migrations must stay backwards compatible with the expand and contract pattern, exactly as with Helm hooks.

What does it cost to adopt GitOps?

Another system to run and upgrade (Argo CD or Flux), a second repo per team, and new failure modes where the controller and git disagree for reasons unrelated to your change, such as mutating webhooks or controllers owning fields. Feedback to the pipeline becomes indirect. And people need to change habits: fixes go through git, not kubectl. Most platform teams find the audit, security and recovery benefits worth it.

In an interview Mid

What problems does GitOps solve compared with running kubectl or helm from a pipeline?

Three problems of push-based CD:

Plus a free audit trail: promotion is a pull request changing one digest line, the review is the approval, the merge commit is the record, and rollback is git revert.

How it is laid out: an app repo (code, Dockerfile, pipeline) and a config repo (what runs where), with environments as directories (overlays/dev, overlays/prod) on main, not branches, and CODEOWNERS on prod.

Also asked: How would you structure a GitOps config repository for several services and environments? · Why should environments be directories rather than branches in a config repo? · Someone hotfixed production with kubectl edit and a minute later the change was gone. Explain.

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.