OnCallReady

Lesson 14.15 · Terraform in Real Life & the Associate Exam · 27 min read

The pipeline: plan as an artifact, gates, approval, apply

In plain words

When a school trip needs a parent's permission, the teacher sends home a specific form: "Tuesday, the museum, back by 3". The parent signs that form. If the trip changes to Wednesday and the zoo, the teacher cannot use the old signature; a new form must be signed. And only the school bus takes the kids, never a random car.

A Terraform pipeline works the same way. terraform plan -out=tfplan is the form, stored as an artifact. A human approves that specific plan at a gate. terraform apply tfplan does exactly what was signed, and refuses with "Saved plan is stale" if the state changed in between. The pipeline, authenticated with OIDC, is the only bus to production.

The problem: "who applied that, and what did they apply?"

On a team, Terraform should not be run from people's laptops against production. Laptops have different Terraform versions, different logins, and no record of what was reviewed. The classic outage: someone reviews a plan on Friday, and on Monday a different plan gets applied, because the code or the cloud changed in between.

The fix is a pipeline: a script that a server runs automatically on every push or pull request (PR). In the jargon this is CI (continuous integration: check every change automatically) and CD (continuous delivery: deploy it the same automatic way). For Terraform, the pipeline runs the checks from 14.6 and 14.10, makes a plan, waits for a human to approve that plan, and applies exactly that plan. Nobody else can apply to production.

What you need to know already: terraform plan -out, show -json, and -detailed-exitcode (12.24, 12.26), state locking and -lock-timeout (13.11, 13.13), prevent_destroy (13.46), jq select and test (7.11, 7.13), set -euo pipefail (6.1), Azure login methods incl. OIDC (12.3), the checks from 14.6 and 14.10.

Words for this lesson

The shape of a Terraform pipeline

on pull request                          on merge to main (per environment)
  fmt -check                               init -input=false
  init -backend=false                      plan -out=tfplan -lock-timeout=10m
  validate                                 show -json tfplan > tfplan.json
  tflint                                   checkov -f tfplan.json
  checkov -d .                             policy gate (jq over tfplan.json)
  plan (per env) -> comment on the PR      publish tfplan + tfplan.json as ARTIFACTS
                                           ---- MANUAL APPROVAL ----
                                           apply tfplan  (the same artifact)

Left column: what runs on every PR - the cheap checks, then a plan per environment posted on the PR so reviewers see it. Right column: what runs once the PR is merged into main - a fresh plan, scans and a gate on it, the plan saved as an artifact, a human approval, then the apply of that saved file.

Three ideas carry all of it:

  1. Plan once, apply that plan. The file reviewed is the file applied.
  2. A human approves a specific plan, not "whatever main produces".
  3. The pipeline is the only path to production. No laptop applies.

Plan as an artifact

terraform plan -input=false -lock-timeout=10m -out=tfplan
terraform show -no-color tfplan > tfplan.txt     # what the approver reads
terraform show -json tfplan > tfplan.json        # what the gates read
flagmeans
-input=falsenever stop to ask a question; fail instead (a pipeline has nobody to answer)
-lock-timeout=10mif the state is locked by another run, wait up to 10 minutes for it (13.13)
-out=tfplansave the plan to the file tfplan
show -no-colorprint the saved plan as text, without colour codes (they look like garbage in a file)
show -jsonprint the saved plan as JSON, for tools

The saved plan remembers the state serial (the version number Terraform bumps on every state write, 13.1) it was computed against. If anything writes that state between plan and apply - another run, a colleague - the apply refuses:

│ Error: Saved plan is stale
│
│ The given plan file can no longer be applied because the state was changed by
│ another operation after the plan was created.

That refusal is a feature: the approval was for a world that no longer exists. Make a new plan and get it approved again.

Treat plan files as secrets: they contain every value Terraform knows, including sensitive ones, in plain text. Keep artifacts for a short time only (retention: how long the server keeps them) and restrict who can download them.

Approval gates

Pipeline services have a built-in way to pause for a human. In GitHub Actions (GitHub's pipeline service, configured in YAML files in the repo) it is an environment with required reviewers. In Azure DevOps (Microsoft's equivalent) it is an environment with an approval check. (These "environments" are pipeline settings named after dev/prod, not your Terraform folders.)

A sketch of the GitHub version - you do not need to write this from memory, only to read it:

# GitHub Actions (sketch)
jobs:
  plan:
    runs-on: ubuntu-latest
    permissions: { id-token: write, contents: read }
    steps:
      - uses: actions/checkout@v4
      - uses: hashicorp/setup-terraform@v3
        with: { terraform_version: 1.9.8 }
      - run: terraform -chdir=envs/prod init -input=false
      - run: terraform -chdir=envs/prod plan -input=false -out=tfplan
      - uses: actions/upload-artifact@v4
        with: { name: tfplan-prod, path: envs/prod/tfplan }

  apply:
    needs: plan
    environment: production          # required reviewers approve here
    concurrency: { group: tf-prod, cancel-in-progress: false }
    steps:
      - uses: actions/download-artifact@v4
        with: { name: tfplan-prod, path: envs/prod }
      - run: terraform -chdir=envs/prod init -input=false
      - run: terraform -chdir=envs/prod apply -input=false tfplan

Line by line, roughly:

Details that matter:

Policy gates with jq

A policy gate is a script that reads the plan JSON and fails the pipeline if the plan does something forbidden. The plan JSON from terraform show -json has a stable format (format_version 1.x). Its list resource_changes has one entry per resource, each with an address, a type and change.actions:

["create"]            new
["update"]            in place
["delete"]            destroy
["delete","create"]   replace (destroy first)
["create","delete"]   replace (create_before_destroy)
["no-op"]             unchanged
["forget"]            removed block with destroy = false (1.7+)

A replacement shows up as two actions - so a check for "delete" catches replacements too, which is what you want: to the data inside a storage account, a replacement is a deletion.

Questions a gate asks, as jq:

# every address that will be destroyed (including replacements)
jq -r '.resource_changes[] | select(.change.actions | index("delete")) | .address' tfplan.json

# how many deletes
jq '[.resource_changes[] | select(.change.actions | index("delete"))] | length' tfplan.json

# deletes of stateful types - block these without an extra approval
jq -r '.resource_changes[]
  | select(.change.actions | index("delete"))
  | select(.type | test("storage_account|key_vault|mssql|postgresql|kubernetes_cluster"))
  | .address' tfplan.json

# why each replacement happens
jq -r '.resource_changes[] | select(.action_reason) | "\(.address)\t\(.action_reason)"' tfplan.json

Reading the third one: go through every resource change; keep those whose actions include "delete"; keep those whose type matches a regular expression (7.3) of stateful types - resources that hold data you cannot get back (storage, vaults, databases, clusters); print their addresses. -r prints raw strings without quotes.

(The lab's jq has no index/1 on arrays; select(.change.actions[] == "delete") does the same job there - simulator.)

A gate script:

#!/usr/bin/env bash
set -euo pipefail
plan_json="$1"
protected=$(jq -r '.resource_changes[]
  | select(.change.actions[] == "delete")
  | select(.type | test("storage_account|key_vault|kubernetes_cluster"))
  | .address' "$plan_json")
if [[ -n "$protected" ]]; then
  echo "BLOCKED: this plan destroys stateful resources:"
  echo "$protected"
  exit 1
fi
echo "gate: ok"

It takes the JSON file as its first argument ($1), collects the protected addresses, and exits 1 with the list if there are any (-n = "string is not empty", 6.10). It turns "somebody should have read the plan" into a check that cannot be skipped. Pair it with prevent_destroy on the resources themselves - two independent layers.

Exit codes the pipeline relies on

terraform fmt -check            0 ok, 3 needs formatting
terraform validate              0 ok, 1 invalid
terraform plan                  0 ok (changes or not), 1 error
terraform plan -detailed-exitcode   0 no changes, 1 error, 2 changes
tflint                          0 clean, 2 issues, 1 tflint error
checkov                         0 clean (or --soft-fail), 1 failed checks

plan -detailed-exitcode is also how a pipeline skips the approval stage when there is nothing to apply (exit 0), and how the nightly drift job decides whether to raise an alarm (exit 2; you build it in 14.19).

Running it locally the same way

The lab's pipeline is a set of shell scripts - there is no CI server in the lab (simulator) - but the steps are the ones a real runner executes. Keeping the steps in scripts (or a Makefile, a file of named commands run with make) that the pipeline calls means developers can run exactly the same checks before pushing.

Anti-patterns to name in an interview

The same pipeline in Azure DevOps

Many companies, especially Microsoft-heavy ones, use Azure DevOps instead of GitHub Actions. The pieces map one to one:

# azure-pipelines.yml (sketch)
stages:
  - stage: plan_prod
    jobs:
      - job: plan
        steps:
          - task: TerraformInstaller@1
            inputs: { terraformVersion: 1.9.8 }
          - task: AzureCLI@2               # service connection with workload identity federation
            inputs:
              azureSubscription: sc-platform-prod-plan
              scriptType: bash
              addSpnToEnvironment: true
              scriptLocation: inlineScript
              inlineScript: ./ci/plan.sh
          - publish: envs/prod/artifacts
            artifact: tfplan-prod

  - stage: apply_prod
    dependsOn: plan_prod
    jobs:
      - deployment: apply
        environment: platform-prod          # approvals and checks are configured here
        strategy:
          runOnce:
            deploy:
              steps:
                - download: current
                  artifact: tfplan-prod
                - script: ./ci/apply.sh

The same shape: a plan stage that installs a pinned Terraform, logs in to Azure through a service connection (Azure DevOps' stored login - here using workload identity federation, its name for the OIDC exchange above), runs ./ci/plan.sh and publishes the artifacts; then an apply stage bound to an environment with an approval, which downloads the artifacts and runs ./ci/apply.sh.

GitHub Actions                     Azure DevOps
environment: production            environment with Approvals and checks
upload/download-artifact           publish / download pipeline artifacts
id-token: write + azure/login      service connection, workload identity federation
concurrency group                  exclusive lock check on the environment

Task names and inputs vary with versions - treat the YAML as the shape, not a copy-paste target.

Identity: what each stage may do

stage              identity                         rights
checks (PR)        none                             no cloud access at all
plan               plan identity per environment    Reader on the scope, read+lease on its state blob
apply              apply identity per environment   Contributor (or narrower) on the scope, write on its state
drift (nightly)    plan identity                    as plan, with -lock=false

(Reader and Contributor are built-in Azure roles: read everything, and change everything except permissions. A lease is how the azurerm backend takes the state lock, 13.11.)

Artifact hygiene

A saved plan contains every value Terraform knows - including secrets in plain text. Treat plan artifacts like state:

# what an approver reads
terraform show -no-color artifacts/tfplan > artifacts/tfplan.txt
# what policy reads
terraform show -json artifacts/tfplan > artifacts/tfplan.json
# a one-line summary for the PR comment
jq -r '[.resource_changes[] | .change.actions | join("/")] | group_by(.) | map("\(.[0]): \(length)") | join(", ")' artifacts/tfplan.json

The last command joins each resource's actions into one word (create, delete/create...), groups equal words, and prints a count per group - e.g. create: 2, update: 1.

What you can now do:

Why it helps

This is the part of Terraform you touch every day in a platform team. Situations: an approver asks why the apply stage failed as stale; someone else's run changed state between plan and apply, and the right answer is re-plan and re-approve. A jq gate blocks a PR because the plan deletes a Key Vault, catching what a tired reviewer missed at the summary line. A security audit asks what the pipeline identity can do; separate plan and apply identities per environment with OIDC is the answer they want. You will also be asked in interviews to design exactly this pipeline, and to name the anti-patterns: auto-approve, re-planning at apply, cancelling a running apply.

FAQ

Why not just run terraform apply -auto-approve after merge?

Because nobody reviewed that plan. A bare apply re-plans against current state, code and data sources, so it can differ from anything seen in the PR. Saving the plan, having a human approve that artifact, and applying the file guarantees the reviewed change is the applied change, and a stale plan fails rather than applying something new.

Why can the plan stage not use the same identity as the apply stage?

Least privilege. The plan identity needs to read everything it manages and to take the state lock (a lease is a write on the blob), but should not be able to change resources. The apply identity needs Contributor, or narrower, on the scope. Separating them, per environment, means a compromised PR pipeline cannot change production, and only the protected branch's pipeline can use the apply identity.

What goes into the plan artifacts, and why are they sensitive?

Three renderings: the binary tfplan for the apply stage, terraform show -no-color tfplan as text for the approver, and terraform show -json tfplan for policy gates. The plan contains every value Terraform knows, including secrets in plaintext. So artifacts get short retention (days), access for the pipeline and approvers only, and are never attached to tickets or chat.

How do I write a gate that blocks destroys?

Read resource_changes from the plan JSON and select entries whose change.actions contain "delete". Replacements show as ["delete","create"] or ["create","delete"], so they count too. Filter on stateful types (storage accounts, key vaults, clusters, databases) and exit 1 if any are found. Pair it with prevent_destroy on those resources: two independent layers.

How does -detailed-exitcode help the pipeline?

terraform plan -detailed-exitcode exits 0 for no changes, 1 for an error and 2 when there are changes. A pipeline uses it to skip the approval and apply stages when there is nothing to do, and the nightly drift job uses exit 2 to open a ticket. Plain plan exits 0 whether or not there are changes.

In an interview Mid

What does a good CI setup for Terraform look like?

On every pull request: fmt -check, validate, tflint, checkov, and a plan per environment posted on the PR for reviewers.

On merge to main, per environment:

  1. terraform plan -input=false -lock-timeout=10m -out=tfplan - the plan is saved as an artifact, with a text version for the approver and JSON for tools.
  2. Gates on that exact plan: a checkov plan scan, and a jq policy gate that fails if resource_changes would delete anything stateful.
  3. A human approval of that specific plan.
  4. terraform apply tfplan - the file reviewed is the file applied. If state changed in between, it fails with Saved plan is stale: re-plan and re-approve, never switch to -auto-approve.

Around it: login with OIDC (no stored secret), separate identities for plan and apply and per environment, one run per state at a time, never cancel an apply halfway, short retention on plan artifacts (they contain secrets). The pipeline is the only path to prod.

Also asked: What does "plan as an artifact" mean? · How do you authenticate a Terraform pipeline without a stored secret? · The apply fails with "Saved plan is stale". What happened, and what do you do?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.