The problem
In chapter 25 the pipeline ran helm upgrade against the cluster. That works, but three things hurt. Every pipeline (and every agent that runs one) holds credentials that can change prod. The cluster is simply whatever the last run made it. And anyone with kubectl can quietly change prod, and nobody notices until the next deploy overwrites it - or does not. GitOps is the fix for all three.
What you need to know already: 25.1 (delivery, CI and CD), 25.2 (git: commits, branches, pull requests, revert), 25.11 (promotion by digest), 25.19 (Helm), 15.9 (the reconciliation loop), 15.40 (declarative YAML and kubectl apply).
Push and pull
What chapter 25 did is push-based CD: something outside the cluster (the pipeline) pushes changes in. It needs cluster credentials, and the cluster's state is the sum of whatever runs happened.
GitOps turns it around: git holds the desired state, and a program inside the cluster pulls it and applies it. The desired state is what you want running (the YAML in git); the live state is what actually runs in the cluster.
The four principles
The OpenGitOps project (a group that wrote down what "GitOps" means) defines four principles:
- Declarative. The desired state is written as data (YAML, a Helm chart and values, a kustomize overlay - 26.2), not as a list of commands to run.
- Versioned and immutable. That data is stored somewhere that keeps every version and never rewrites old ones: git. Every change has an author, a review, a timestamp and a revert.
- Pulled automatically. Software agents in (or next to) the cluster pull the desired state. Nothing outside pushes into the cluster.
- Continuously reconciled. The agents keep comparing live state with desired state and fix the difference - not only at deploy time, always. This is the reconciliation loop from 15.9, with git as the source.
When live and desired differ, that difference is called drift. Fixing it is reconciling.
The agent is Argo CD (lesson 26.4) or a similar tool called Flux. The pipeline's job shrinks: build, test, push an image, and commit the new image digest to the config repo. Argo CD notices the commit and applies it.
Why it is worth it
push (ch25) pull (GitOps)
who has cluster creds every pipeline, every agent only the in-cluster controller
source of truth the cluster + the last run git
drift silent until the next deploy detected in minutes, optionally reverted
audit pipeline logs (kept N days) git log / pull request history
rollback re-run an old pipeline git revert (or argocd rollback)
disaster recovery re-run every pipeline, in order point a new cluster at the repo
The credentials line is the one security teams care about. In a bank, "the CI system can kubectl apply anything to prod" is an audit finding; "a controller in the cluster applies what was merged to a protected branch, after review" is a control (a check the auditors accept).
What it costs: another moving part to run and understand, a second repo, and a new way to fail - the controller and git can disagree for reasons that have nothing to do with your change. You meet two of them as incidents in this chapter.
Two repos
pay/payments-api the APP repo: code, Dockerfile, tests, the pipeline
-> produces images (and maybe the chart)
pay/payments-gitops the CONFIG repo: what runs where - base + overlays / values per env
-> read by Argo CD
Why split them:
- Different lifecycles. A config change (replicas, a timeout, a new environment) should not rebuild and retest the app; an app change should not need a config review.
- Different permissions. Many developers push to the app repo; few may approve prod config. Branch protection (25.2) and a
CODEOWNERSfile (a file in the repo that names who must approve changes to which paths) onoverlays/prod/do that. - No loops. If the pipeline commits the new digest into the app repo, that commit triggers the pipeline again. Into the config repo it triggers nothing but Argo CD.
- A readable history.
git log overlays/prodis the list of prod deployments.
Environments are directories, not branches
base/ the service, the same everywhere
overlays/dev/ what differs in dev
overlays/prod/ what differs in prod
Branch-per-environment (dev, staging, prod branches) looks natural and fails in practice: promoting becomes a merge that drags unrelated changes along, branches drift apart, and "what is different between staging and prod" becomes a merge puzzle instead of diff -r overlays/staging overlays/prod (diff -r compares two directories file by file). Every Argo CD Application points at the main branch, each at a different path.
Promotion and rollback
Promotion to prod (25.11: moving the same built image to the next environment) is a pull request that changes one line:
images:
- name: registry.lab/pay/payments-api
- digest: sha256:4be1...
+ digest: sha256:9c3e...
The approval is the pull request review (plus branch protection), the audit record is the merge commit, the deploy is Argo CD noticing it. Rollback is git revert of that commit (25.2) - a new commit, reviewed like any other. Tools such as Argo CD Image Updater, Flux image automation or Renovate can open these pull requests for you; for dev many teams let the pipeline commit directly.
What GitOps does not solve
- Secrets. You cannot commit passwords to git - lesson 26.22.
- Database migrations. Ordering and compatibility are still yours.
- The build itself. That stays in CI.
- kubectl access. It does not go away by itself: humans keep read access, write access is removed, and break-glass (an emergency procedure to get write access, logged and reviewed) is written down. An incident later in this chapter is what happens when someone fixes prod with kubectl anyway.
What you can now do
- Explain push-based CD vs GitOps and the four principles in your own words.
- Say why the config lives in its own repo, with environments as directories.
- Describe a promotion (a commit that changes a digest) and a rollback (
git revert).