The problem: "who applied that, and what did they apply?"
On a team, Terraform should not be run from people's laptops against production. Laptops have different Terraform versions, different logins, and no record of what was reviewed. The classic outage: someone reviews a plan on Friday, and on Monday a different plan gets applied, because the code or the cloud changed in between.
The fix is a pipeline: a script that a server runs automatically on every push or pull request (PR). In the jargon this is CI (continuous integration: check every change automatically) and CD (continuous delivery: deploy it the same automatic way). For Terraform, the pipeline runs the checks from 14.6 and 14.10, makes a plan, waits for a human to approve that plan, and applies exactly that plan. Nobody else can apply to production.
What you need to know already: terraform plan -out, show -json, and -detailed-exitcode (12.24, 12.26), state locking and -lock-timeout (13.11, 13.13), prevent_destroy (13.46), jq select and test (7.11, 7.13), set -euo pipefail (6.1), Azure login methods incl. OIDC (12.3), the checks from 14.6 and 14.10.
Words for this lesson
- Runner (or agent): the machine the pipeline's commands run on - usually a fresh Linux VM or container (10.3) that exists only for that run.
- Stage / job / step: a pipeline is split into stages (plan, apply), each made of jobs, each made of steps (single commands).
- Artifact: a file one stage saves so a later stage can download it - here, the saved plan.
- Approval gate: a point where the pipeline pauses until a named person clicks "approve".
- Gate in general: any check that can stop the pipeline.
The shape of a Terraform pipeline
on pull request on merge to main (per environment)
fmt -check init -input=false
init -backend=false plan -out=tfplan -lock-timeout=10m
validate show -json tfplan > tfplan.json
tflint checkov -f tfplan.json
checkov -d . policy gate (jq over tfplan.json)
plan (per env) -> comment on the PR publish tfplan + tfplan.json as ARTIFACTS
---- MANUAL APPROVAL ----
apply tfplan (the same artifact)
Left column: what runs on every PR - the cheap checks, then a plan per environment posted on the PR so reviewers see it. Right column: what runs once the PR is merged into main - a fresh plan, scans and a gate on it, the plan saved as an artifact, a human approval, then the apply of that saved file.
Three ideas carry all of it:
- Plan once, apply that plan. The file reviewed is the file applied.
- A human approves a specific plan, not "whatever main produces".
- The pipeline is the only path to production. No laptop applies.
Plan as an artifact
terraform plan -input=false -lock-timeout=10m -out=tfplan
terraform show -no-color tfplan > tfplan.txt # what the approver reads
terraform show -json tfplan > tfplan.json # what the gates read
| flag | means |
|---|---|
-input=false | never stop to ask a question; fail instead (a pipeline has nobody to answer) |
-lock-timeout=10m | if the state is locked by another run, wait up to 10 minutes for it (13.13) |
-out=tfplan | save the plan to the file tfplan |
show -no-color | print the saved plan as text, without colour codes (they look like garbage in a file) |
show -json | print the saved plan as JSON, for tools |
The saved plan remembers the state serial (the version number Terraform bumps on every state write, 13.1) it was computed against. If anything writes that state between plan and apply - another run, a colleague - the apply refuses:
│ Error: Saved plan is stale
│
│ The given plan file can no longer be applied because the state was changed by
│ another operation after the plan was created.
That refusal is a feature: the approval was for a world that no longer exists. Make a new plan and get it approved again.
Treat plan files as secrets: they contain every value Terraform knows, including sensitive ones, in plain text. Keep artifacts for a short time only (retention: how long the server keeps them) and restrict who can download them.
Approval gates
Pipeline services have a built-in way to pause for a human. In GitHub Actions (GitHub's pipeline service, configured in YAML files in the repo) it is an environment with required reviewers. In Azure DevOps (Microsoft's equivalent) it is an environment with an approval check. (These "environments" are pipeline settings named after dev/prod, not your Terraform folders.)
A sketch of the GitHub version - you do not need to write this from memory, only to read it:
# GitHub Actions (sketch)
jobs:
plan:
runs-on: ubuntu-latest
permissions: { id-token: write, contents: read }
steps:
- uses: actions/checkout@v4
- uses: hashicorp/setup-terraform@v3
with: { terraform_version: 1.9.8 }
- run: terraform -chdir=envs/prod init -input=false
- run: terraform -chdir=envs/prod plan -input=false -out=tfplan
- uses: actions/upload-artifact@v4
with: { name: tfplan-prod, path: envs/prod/tfplan }
apply:
needs: plan
environment: production # required reviewers approve here
concurrency: { group: tf-prod, cancel-in-progress: false }
steps:
- uses: actions/download-artifact@v4
with: { name: tfplan-prod, path: envs/prod }
- run: terraform -chdir=envs/prod init -input=false
- run: terraform -chdir=envs/prod apply -input=false tfplan
Line by line, roughly:
- Two jobs,
planandapply.needs: planmakesapplywait forplan. runs-on: ubuntu-latest- run on a fresh Ubuntu runner.uses: ...steps are ready-made actions: check out the repo, install a pinned Terraform, upload / download an artifact.run: terraform -chdir=envs/prod ...--chdir=DIRmakes Terraform act as if you hadcd'd intoDIRfirst.environment: production- this is where the job pauses for approval.concurrency- only one job in the grouptf-prodat a time.
Details that matter:
- The apply job downloads the artifact; it never re-plans.
concurrencywithcancel-in-progress: false: two applies must never overlap, and a running apply must never be cancelled halfway (that leaves a held lock and half-created resources, 13.12).-input=falseeverywhere: a pipeline must fail, not hang on a prompt.- Authentication with OIDC (12.3):
id-token: writelets the job get a short-lived token from GitHub, which Azure is set up to trust and swap for an Azure login. Terraform readsARM_USE_OIDC=trueplusARM_CLIENT_ID,ARM_TENANT_IDandARM_SUBSCRIPTION_ID(which login, in which directory, for which subscription - none of them secret). No password exists to leak or rotate. - Separate identities per environment: the dev pipeline's login has no rights in prod, not even on prod's state.
Policy gates with jq
A policy gate is a script that reads the plan JSON and fails the pipeline if the plan does something forbidden. The plan JSON from terraform show -json has a stable format (format_version 1.x). Its list resource_changes has one entry per resource, each with an address, a type and change.actions:
["create"] new
["update"] in place
["delete"] destroy
["delete","create"] replace (destroy first)
["create","delete"] replace (create_before_destroy)
["no-op"] unchanged
["forget"] removed block with destroy = false (1.7+)
A replacement shows up as two actions - so a check for "delete" catches replacements too, which is what you want: to the data inside a storage account, a replacement is a deletion.
Questions a gate asks, as jq:
# every address that will be destroyed (including replacements)
jq -r '.resource_changes[] | select(.change.actions | index("delete")) | .address' tfplan.json
# how many deletes
jq '[.resource_changes[] | select(.change.actions | index("delete"))] | length' tfplan.json
# deletes of stateful types - block these without an extra approval
jq -r '.resource_changes[]
| select(.change.actions | index("delete"))
| select(.type | test("storage_account|key_vault|mssql|postgresql|kubernetes_cluster"))
| .address' tfplan.json
# why each replacement happens
jq -r '.resource_changes[] | select(.action_reason) | "\(.address)\t\(.action_reason)"' tfplan.json
Reading the third one: go through every resource change; keep those whose actions include "delete"; keep those whose type matches a regular expression (7.3) of stateful types - resources that hold data you cannot get back (storage, vaults, databases, clusters); print their addresses. -r prints raw strings without quotes.
(The lab's jq has no index/1 on arrays; select(.change.actions[] == "delete") does the same job there - simulator.)
A gate script:
#!/usr/bin/env bash
set -euo pipefail
plan_json="$1"
protected=$(jq -r '.resource_changes[]
| select(.change.actions[] == "delete")
| select(.type | test("storage_account|key_vault|kubernetes_cluster"))
| .address' "$plan_json")
if [[ -n "$protected" ]]; then
echo "BLOCKED: this plan destroys stateful resources:"
echo "$protected"
exit 1
fi
echo "gate: ok"
It takes the JSON file as its first argument ($1), collects the protected addresses, and exits 1 with the list if there are any (-n = "string is not empty", 6.10). It turns "somebody should have read the plan" into a check that cannot be skipped. Pair it with prevent_destroy on the resources themselves - two independent layers.
Exit codes the pipeline relies on
terraform fmt -check 0 ok, 3 needs formatting
terraform validate 0 ok, 1 invalid
terraform plan 0 ok (changes or not), 1 error
terraform plan -detailed-exitcode 0 no changes, 1 error, 2 changes
tflint 0 clean, 2 issues, 1 tflint error
checkov 0 clean (or --soft-fail), 1 failed checks
plan -detailed-exitcode is also how a pipeline skips the approval stage when there is nothing to apply (exit 0), and how the nightly drift job decides whether to raise an alarm (exit 2; you build it in 14.19).
Running it locally the same way
The lab's pipeline is a set of shell scripts - there is no CI server in the lab (simulator) - but the steps are the ones a real runner executes. Keeping the steps in scripts (or a Makefile, a file of named commands run with make) that the pipeline calls means developers can run exactly the same checks before pushing.
Anti-patterns to name in an interview
terraform apply -auto-approvestraight from a merge, with no saved plan.- One pipeline identity with Owner (full control) on every subscription.
- Plans reviewed as a summary line ("3 to add, 1 to change") without the body.
- Approvals given on a PR plan, then the apply stage makes a new plan.
- Cancelling a stuck apply instead of investigating the lock.
- Secrets passed as
-var db_password=...in the pipeline: they end up in state anyway, and often in logs. Lesson 14.20.
The same pipeline in Azure DevOps
Many companies, especially Microsoft-heavy ones, use Azure DevOps instead of GitHub Actions. The pieces map one to one:
# azure-pipelines.yml (sketch)
stages:
- stage: plan_prod
jobs:
- job: plan
steps:
- task: TerraformInstaller@1
inputs: { terraformVersion: 1.9.8 }
- task: AzureCLI@2 # service connection with workload identity federation
inputs:
azureSubscription: sc-platform-prod-plan
scriptType: bash
addSpnToEnvironment: true
scriptLocation: inlineScript
inlineScript: ./ci/plan.sh
- publish: envs/prod/artifacts
artifact: tfplan-prod
- stage: apply_prod
dependsOn: plan_prod
jobs:
- deployment: apply
environment: platform-prod # approvals and checks are configured here
strategy:
runOnce:
deploy:
steps:
- download: current
artifact: tfplan-prod
- script: ./ci/apply.sh
The same shape: a plan stage that installs a pinned Terraform, logs in to Azure through a service connection (Azure DevOps' stored login - here using workload identity federation, its name for the OIDC exchange above), runs ./ci/plan.sh and publishes the artifacts; then an apply stage bound to an environment with an approval, which downloads the artifacts and runs ./ci/apply.sh.
GitHub Actions Azure DevOps
environment: production environment with Approvals and checks
upload/download-artifact publish / download pipeline artifacts
id-token: write + azure/login service connection, workload identity federation
concurrency group exclusive lock check on the environment
Task names and inputs vary with versions - treat the YAML as the shape, not a copy-paste target.
Identity: what each stage may do
stage identity rights
checks (PR) none no cloud access at all
plan plan identity per environment Reader on the scope, read+lease on its state blob
apply apply identity per environment Contributor (or narrower) on the scope, write on its state
drift (nightly) plan identity as plan, with -lock=false
(Reader and Contributor are built-in Azure roles: read everything, and change everything except permissions. A lease is how the azurerm backend takes the state lock, 13.11.)
- The plan stage must read everything it manages and lock the state (a lease counts as a write on the state file), but it should not be able to change resources.
- Only the apply identity can change prod, and only the pipeline on the protected
mainbranch can use it. - No human has standing write access to prod state. Break-glass access (an emergency login, used only when the normal path is broken) exists, is logged, and is the exception - you meet its consequences in the stale-plan incident, 14.18.
Artifact hygiene
A saved plan contains every value Terraform knows - including secrets in plain text. Treat plan artifacts like state:
- short retention (days, not months), readable only by the pipeline and the approvers;
- publish the text rendering (
show -no-color) for humans and the JSON for machines, and keep the binary plan only for the apply stage; - never attach plans to tickets or chat.
# what an approver reads
terraform show -no-color artifacts/tfplan > artifacts/tfplan.txt
# what policy reads
terraform show -json artifacts/tfplan > artifacts/tfplan.json
# a one-line summary for the PR comment
jq -r '[.resource_changes[] | .change.actions | join("/")] | group_by(.) | map("\(.[0]): \(length)") | join(", ")' artifacts/tfplan.json
The last command joins each resource's actions into one word (create, delete/create...), groups equal words, and prints a count per group - e.g. create: 2, update: 1.
What you can now do:
- Describe a Terraform pipeline from PR to apply, and why the apply uses the saved plan.
- Write a jq policy gate over
resource_changesthat blocks destroys. - Explain OIDC for pipelines and why plan and apply get different identities.