A new joiner runs their first plan for the checkout dev environment, following the wiki: copy main.tf from the prod folder and change the names to dev. They have not applied yet, and they ask whether this is normal:
$ terraform plan
...
# azurerm_resource_group.main must be replaced
-/+ resource "azurerm_resource_group" "main" {
~ id = "/subscriptions/00000000-1111-2222-3333-444444444444/resourceGroups/rg-checkout-prod" -> (known after apply)
~ name = "rg-checkout-prod" -> "rg-checkout-dev" # forces replacement
}
# azurerm_virtual_network.main must be replaced
-/+ resource "azurerm_virtual_network" "main" {
~ name = "vnet-checkout-prod" -> "vnet-checkout-dev" # forces replacement
~ resource_group_name = "rg-checkout-prod" -> "rg-checkout-dev" # forces replacement
...
}
Plan: 2 to add, 0 to change, 2 to destroy.It is not normal. Applied, this plan deletes production's resource group (and everything in it) and builds dev-named copies.
What is happening: the plan is honest, the state is wrong
Terraform compares three things: your configuration (the .tf files), its state (the record of which real objects it manages, with their IDs and attributes), and the real infrastructure. A plan is the difference.
The configuration says rg-checkout-dev. If the plan says it will change rg-checkout-prod into it, then the state it read lists prod's resource group. The question is not "what is wrong with these resources" but "which state did this plan read?"
Where the state lives is set by the backend. For Azure that is a blob in a storage account, identified by resource_group_name, storage_account_name, container_name and key. Copy a prod folder and you copy prod's backend block with it.
The diagnosis path
1. Read the plan for prod names on the left side
terraform plan | grep -E "must be|forces replacement|Plan:"must be replaced and # forces replacement mean destroy and create, not an edit. A value on the left of -> is what state holds now. "rg-checkout-prod" -> in a dev plan is the proof that the state is prod's.
2. Look at the backend block
$ grep -A5 'backend "azurerm"' main.tf
backend "azurerm" {
resource_group_name = "rg-tfstate"
storage_account_name = "sttfstatesysop"
container_name = "tfstate"
key = "checkout/prod.tfstate"
}The names were changed to dev. The key was not.
3. Check what init actually recorded
terraform init writes the backend settings it used into .terraform/terraform.tfstate. That file, not main.tf, is what later commands use until you re-init:
$ jq .backend.config.key .terraform/terraform.tfstate
"checkout/prod.tfstate"4. Rule out the other two mix-ups
- Workspaces.
terraform workspace showprints the active one. With CLI workspaces, each workspace is a separate state under the same backend (for azurerm, the blob name gets anenv:<workspace>suffix). Running indefaultwhen you meantdevgives the same symptom. - Variable files. A plan with prod's
-var-fileagainst dev's state shows the opposite: dev names on the left, prod names on the right.
$ terraform workspace show
default
$ terraform state show azurerm_resource_group.main
# azurerm_resource_group.main:
resource "azurerm_resource_group" "main" {
id = "/subscriptions/00000000-1111-2222-3333-444444444444/resourceGroups/rg-checkout-prod"
location = "westeurope"
name = "rg-checkout-prod"
}state show settles it: the state this directory uses holds production.
The fix: point at dev's state, without copying prod into it
Fix the key, then re-initialise. Terraform notices the backend changed and refuses to guess:
$ sed -i 's|checkout/prod.tfstate|checkout/dev.tfstate|' main.tf
$ terraform init
Initializing the backend...
╷
│ Error: Backend configuration changed
│
│ A change in the backend configuration has been detected, which may require
│ migrating existing state.
│
│ If you wish to attempt automatic migration of the state, use "terraform init
│ -migrate-state".
│ If you wish to store the current configuration with no changes to the state,
│ use "terraform init -reconfigure".
╵This choice matters:
-migrate-statecopies the state from the old backend to the new one. Here that would copy prod's state into dev's key and poison dev too.-reconfigurejust switches to the new backend and uses whatever state is already there, which is dev's own.
$ terraform init -reconfigure
...
Successfully configured the backend "azurerm"! Terraform will automatically
use this backend unless the backend configuration changes.
$ terraform plan
...
No changes. Your infrastructure matches the configuration.A clean plan: dev's state matches dev's configuration. Nothing in production was touched, because nothing was applied.
Keeping it from coming back
- Never hand-edit keys per environment. Use a directory per environment with a small backend file each (
terraform init -backend-config=backend-dev.hcl), or derive the key in the pipeline. Copy-paste is how the prod key travels. - Separate state storage and permissions per environment. If the dev pipeline identity cannot even read prod's state container, this plan fails with a 403 instead of proposing a disaster.
- Gate destroys and replaces. Save the plan (
terraform plan -out=tfplan), then count deletes interraform show -json tfplan:jq '[.resource_changes[] | select(.change.actions | index("delete"))] | length'. Anything above zero needs an extra approval. - Read the left side of
->. Whenever a plan touches something you did not mean to change, stop and ask which state you are looking at.
The wrong state is one of the classic Terraform outages. The other two are a count index shift that replaces half a list and two runs writing the same state, which is what the state lock prevents. If a plan like this one ever reaches a pipeline log, also check what else the log holds: secrets in CI output.
Practise it
Incident: the dev plan wants to replace production (13.9) is this directory, copied from prod, with the plan waiting to be read.