OnCallReady

TerraformAzureCI/CDSRE · 4 min read

Terraform plan in dev wants to replace production: wrong state, backend key or workspace

A dev plan that renames prod resources is reading prod's state. How to read the plan, find the backend key, and re-init with -reconfigure, not -migrate-state.

A new joiner runs their first plan for the checkout dev environment, following the wiki: copy main.tf from the prod folder and change the names to dev. They have not applied yet, and they ask whether this is normal:

terminal
$ terraform plan
...
  # azurerm_resource_group.main must be replaced
-/+ resource "azurerm_resource_group" "main" {
      ~ id   = "/subscriptions/00000000-1111-2222-3333-444444444444/resourceGroups/rg-checkout-prod" -> (known after apply)
      ~ name = "rg-checkout-prod" -> "rg-checkout-dev" # forces replacement
    }

  # azurerm_virtual_network.main must be replaced
-/+ resource "azurerm_virtual_network" "main" {
      ~ name                = "vnet-checkout-prod" -> "vnet-checkout-dev" # forces replacement
      ~ resource_group_name = "rg-checkout-prod" -> "rg-checkout-dev" # forces replacement
      ...
    }

Plan: 2 to add, 0 to change, 2 to destroy.

It is not normal. Applied, this plan deletes production's resource group (and everything in it) and builds dev-named copies.

What is happening: the plan is honest, the state is wrong

Terraform compares three things: your configuration (the .tf files), its state (the record of which real objects it manages, with their IDs and attributes), and the real infrastructure. A plan is the difference.

The configuration says rg-checkout-dev. If the plan says it will change rg-checkout-prod into it, then the state it read lists prod's resource group. The question is not "what is wrong with these resources" but "which state did this plan read?"

Where the state lives is set by the backend. For Azure that is a blob in a storage account, identified by resource_group_name, storage_account_name, container_name and key. Copy a prod folder and you copy prod's backend block with it.

The diagnosis path

1. Read the plan for prod names on the left side

output
terraform plan | grep -E "must be|forces replacement|Plan:"

must be replaced and # forces replacement mean destroy and create, not an edit. A value on the left of -> is what state holds now. "rg-checkout-prod" -> in a dev plan is the proof that the state is prod's.

2. Look at the backend block

terminal
$ grep -A5 'backend "azurerm"' main.tf
  backend "azurerm" {
    resource_group_name  = "rg-tfstate"
    storage_account_name = "sttfstatesysop"
    container_name       = "tfstate"
    key                  = "checkout/prod.tfstate"
  }

The names were changed to dev. The key was not.

3. Check what init actually recorded

terraform init writes the backend settings it used into .terraform/terraform.tfstate. That file, not main.tf, is what later commands use until you re-init:

terminal
$ jq .backend.config.key .terraform/terraform.tfstate
"checkout/prod.tfstate"

4. Rule out the other two mix-ups

  • Workspaces. terraform workspace show prints the active one. With CLI workspaces, each workspace is a separate state under the same backend (for azurerm, the blob name gets an env:<workspace> suffix). Running in default when you meant dev gives the same symptom.
  • Variable files. A plan with prod's -var-file against dev's state shows the opposite: dev names on the left, prod names on the right.
terminal
$ terraform workspace show
default
$ terraform state show azurerm_resource_group.main
# azurerm_resource_group.main:
resource "azurerm_resource_group" "main" {
    id       = "/subscriptions/00000000-1111-2222-3333-444444444444/resourceGroups/rg-checkout-prod"
    location = "westeurope"
    name     = "rg-checkout-prod"
}

state show settles it: the state this directory uses holds production.

The fix: point at dev's state, without copying prod into it

Fix the key, then re-initialise. Terraform notices the backend changed and refuses to guess:

terminal
$ sed -i 's|checkout/prod.tfstate|checkout/dev.tfstate|' main.tf
$ terraform init

Initializing the backend...
╷
│ Error: Backend configuration changed
│ 
│ A change in the backend configuration has been detected, which may require
│ migrating existing state.
│ 
│ If you wish to attempt automatic migration of the state, use "terraform init
│ -migrate-state".
│ If you wish to store the current configuration with no changes to the state,
│ use "terraform init -reconfigure".
╵

This choice matters:

  • -migrate-state copies the state from the old backend to the new one. Here that would copy prod's state into dev's key and poison dev too.
  • -reconfigure just switches to the new backend and uses whatever state is already there, which is dev's own.
terminal
$ terraform init -reconfigure
...
Successfully configured the backend "azurerm"! Terraform will automatically
use this backend unless the backend configuration changes.
$ terraform plan
...
No changes. Your infrastructure matches the configuration.

A clean plan: dev's state matches dev's configuration. Nothing in production was touched, because nothing was applied.

Keeping it from coming back

  • Never hand-edit keys per environment. Use a directory per environment with a small backend file each (terraform init -backend-config=backend-dev.hcl), or derive the key in the pipeline. Copy-paste is how the prod key travels.
  • Separate state storage and permissions per environment. If the dev pipeline identity cannot even read prod's state container, this plan fails with a 403 instead of proposing a disaster.
  • Gate destroys and replaces. Save the plan (terraform plan -out=tfplan), then count deletes in terraform show -json tfplan: jq '[.resource_changes[] | select(.change.actions | index("delete"))] | length'. Anything above zero needs an extra approval.
  • Read the left side of ->. Whenever a plan touches something you did not mean to change, stop and ask which state you are looking at.

The wrong state is one of the classic Terraform outages. The other two are a count index shift that replaces half a list and two runs writing the same state, which is what the state lock prevents. If a plan like this one ever reaches a pipeline log, also check what else the log holds: secrets in CI output.

Practise it

Incident: the dev plan wants to replace production (13.9) is this directory, copied from prod, with the plan waiting to be read.

OnCallReady is free, with no ads and no tracking. RSS · All posts