OnCallReady

Lesson 13.4 · Terraform: State & Modules · 28 min read

Backends: where state lives, and how init moves it

In plain words

Your diary is safe in your bedroom only as long as nobody else needs to write in it. If your whole family keeps one shared diary, it should live in a locked cabinet in the hallway where everyone can reach it, with a rule that only one person holds the pen at a time, and a photocopy of every page, in case someone spills juice on it.

A Terraform backend is that cabinet. With no backend block, state is a local file. With backend "azurerm" it is a blob in a storage account, locked with a blob lease. terraform init connects to it and, if you change backends, asks whether to copy your existing state across. Answering "no" to that question is how you end up with an empty cabinet next to a house full of furniture.

The problem

State is a secret, and a whole team needs the same copy of it. A file on your laptop fails both tests: your colleague's plan cannot see it, and anyone who gets your laptop gets every password in it.

So state has to live somewhere shared, protected and lockable. The setting that says where is the backend. This lesson is about choosing one, pointing Terraform at it, and moving existing state into it without losing anything.

What you need to know already:

What a backend is

The backend decides where state is stored and whether it can be locked (locked = only one run may write at a time; 13.11).

In Terraform 1.x a backend is just storage. (Older versions had "enhanced" backends that also ran your commands remotely; today that job belongs to HCP Terraform's cloud block, at the end of this lesson.)

The common ones:

local      a file on disk (the default)             no locking across machines
azurerm    a blob in an Azure Storage container     locking with a blob lease
s3         an object in an S3 bucket                locking with a lock file (1.10+)
gcs        an object in a GCS bucket                locking built in
consul, pg, http ...                                 varies
cloud      HCP Terraform / Terraform Enterprise      runs, locking, history, policy

Words in that table:

The local backend is what you get with no block at all. You can still write it out, for example to keep state outside the code directory:

terraform {
  backend "local" {
    path = "../state/dev.tfstate"
  }
}

The azurerm backend

This is the one the labs use: state as a blob in an Azure storage account.

terraform {
  backend "azurerm" {
    resource_group_name  = "rg-tfstate"
    storage_account_name = "sttfstatesysop"
    container_name       = "tfstate"
    key                  = "platform/dev.tfstate"
    use_azuread_auth     = true
  }
}

Each argument:

resource_group_name   the resource group the storage account lives in
storage_account_name  the storage account (its name is unique across all of Azure)
container_name        the storage container inside it
key                   the blob name - one per state; slashes make it look like a folder
use_azuread_auth      log in to the blob as a person or pipeline identity (Entra ID),
                      not with the account's master key

Two ways Terraform can prove who it is to the storage account:

On a laptop you sign in with the Azure CLI (az login, the az command-line tool). In a pipeline the job signs in with OIDC (the pipeline gets a short-lived token instead of a stored password). The backend uses the same login as the azurerm provider.

Locking is automatic: for every command that could write state, the backend takes a lease on the blob, and releases it when done.

The chicken and the egg

The storage account must exist before terraform init. A backend cannot create its own storage - it needs somewhere to keep the state of the thing that creates the storage.

Options, in order of preference:

  1. The platform team owns a small bootstrap configuration: a separate Terraform configuration (with local state, or state in its own "state of the states" account) that creates the storage accounts, containers, roles, versioning and soft delete once. 13.8 builds one.
  2. A script with the Azure CLI (az storage account create ...), run once and kept in the repo.

Harden that storage account:

No variables in a backend block

It is tempting to build the key from a variable:

terraform {
  backend "azurerm" {
    key = "platform/${var.env}.tfstate"      # not allowed
  }
}
╷
│ Error: Variables not allowed
│
│   on main.tf line 6, in terraform:
│    6:     key                  = var.state_key
│
│ Variables may not be used here.
╵

Why: the backend is set up before variables are evaluated. Terraform needs the state before it can evaluate anything else. The per-environment parts come from partial configuration instead.

Partial configuration: -backend-config

Leave out whatever differs per environment - or the whole body:

terraform {
  backend "azurerm" {}
}

and hand the missing values to init with -backend-config. Either one key=value per flag, or a file:

terraform init -backend-config=backends/dev.hcl
terraform init \
  -backend-config=resource_group_name=rg-tfstate \
  -backend-config=storage_account_name=sttfstatesysop \
  -backend-config=container_name=tfstate \
  -backend-config=key=orders/dev.tfstate

The file holds plain key = value lines, no block around them:

# backends/dev.hcl
resource_group_name  = "rg-tfstate"
storage_account_name = "sttfstatesysop"
container_name       = "tfstate"
key                  = "orders/dev.tfstate"

Pipelines do exactly this: the same code, init with the environment's backend file, then plan. Secrets (an access key, if you really must) come in through environment variables, never in these files.

Changing the backend: init decides what happens to state

Change the backend block or its settings, and every other command refuses until you run init again:

╷
│ Error: Backend initialization required, please run "terraform init"
│
│ Reason: Initial configuration of the requested backend "azurerm"
│ ...
╵

Then init has to know what to do with the state you already have. The flags:

(no flag), local -> azurerm   init offers to COPY the local state into the blob
                              - answer yes, or you start with an empty state
-migrate-state                copy state from the old backend/key to the new one
-reconfigure                  use the new settings and IGNORE the old state entirely
-force-copy                   answer "yes" to the copy question without asking (for CI)

The copy prompt, word for word:

Do you want to copy existing state to the new backend?
  Pre-existing state was found while migrating the previous "local" backend to the
  newly configured "azurerm" backend. No existing state was found in the newly
  configured "azurerm" backend. Do you want to copy this state to the new "azurerm"
  backend? Enter "yes" to copy and "no" to start with an empty state.

Answering no is how teams end up with an empty remote state next to live infrastructure. The next plan thinks nothing exists and wants to create everything, and every create fails with already exists.

-reconfigure vs -migrate-state is the distinction to be precise about:

After a local -> remote move, the old terraform.tfstate is still on disk, secrets and all. Delete it once you have checked the remote copy.

What init recorded

init writes the backend settings it worked out into .terraform/terraform.tfstate. Despite the name this is not your state - just a small note of which backend to use:

# after init with the azurerm backend (the remote-state mission)
jq .backend .terraform/terraform.tfstate
{
  "type": "azurerm",
  "config": {
    "container_name": "tfstate",
    "key": "orders/dev.tfstate",
    "resource_group_name": "rg-tfstate",
    "storage_account_name": "sttfstatesysop"
  },
  "hash": "..."
}

type is the backend kind, config the resolved settings, hash a fingerprint Terraform uses to notice that the block changed.

That file is how a later terraform plan knows which blob to read - and why plan in a freshly cloned directory fails until you run init. It lives in .terraform/, so it is never committed.

Reading remote state

The state commands work against any backend, local or remote:

terraform state list                           # every address in the state
terraform state pull > /tmp/backup.tfstate     # a local copy, e.g. before surgery
terraform output -json                         # the outputs, as JSON

In the lab, the remote blobs live in a simulated storage account. There is no az storage blob command, so terraform state pull is how you look inside (simulator).

Other backends, briefly (exam material)

The same idea on AWS:

terraform {
  backend "s3" {
    bucket       = "acme-tfstate"
    key          = "platform/prod.tfstate"
    region       = "eu-west-1"
    use_lockfile = true           # 1.10+: locking with a lock file next to the state
    encrypt      = true
  }
}

Older S3 setups locked through a separate database table (DynamoDB); that way is now deprecated.

HCP Terraform is HashiCorp's hosted service for running Terraform. It replaces the backend block with a cloud block (Terraform 1.1+):

terraform {
  cloud {
    organization = "acme"
    workspaces {
      name = "platform-prod"
    }
  }
}

With cloud, state lives in HCP Terraform, and terraform plan can run on HashiCorp's machines while streaming the output to you. You also get run history, locking, shared variables, policy checks and a private module registry. A workspace there is one state plus its settings.

Later (Ch 14): HCP Terraform in detail, for the exam.

Checklist for a real state backend

What you can now do:

Why it helps

You will set up or fix backends at least once in every new team. Situations: a pipeline needs the same code against dev and prod, so you use partial configuration with -backend-config=backends/prod.hcl, and must use -reconfigure when switching keys, not -migrate-state, or you copy dev's state over prod's. A colleague migrates local state to Azure, answers "no" to the copy prompt, and the next plan wants to create 80 resources that already exist. A security review flags ARM_ACCESS_KEY in pipelines; you move to use_azuread_auth with Storage Blob Data Contributor scoped to one container. The exam asks why backends cannot use variables and what -migrate-state does.

FAQ

Why can the backend block not use variables?

The backend is configured during init, before Terraform evaluates variables, because Terraform needs to know where state is before it can evaluate anything. So the block accepts only literals. Per-environment values come from partial configuration: leave the block empty or incomplete and pass -backend-config=file.hcl or -backend-config=key=value to init.

What is the difference between -reconfigure and -migrate-state?

-migrate-state copies state from the old backend configuration to the new one: use it when moving a state to a new home, like a new storage account or a renamed key. -reconfigure uses the new configuration and ignores the old state entirely: use it when you deliberately point the same code at a different state, such as a pipeline switching from the dev key to the prod key. Mixing them up can copy one environment's state over another.

Who creates the storage account for the backend?

Not the configuration that uses it: a backend cannot create its own storage. Usually a small bootstrap configuration owned by the platform team (with local state, or state in a separate account) creates the storage accounts, containers, RBAC, versioning, soft delete and a resource lock once. A checked-in az script is the simpler alternative.

Should the backend use an access key or Entra ID?

Entra ID. With use_azuread_auth = true the pipeline authenticates with its identity (OIDC or managed identity) and needs Storage Blob Data Contributor on the container, which you can scope precisely. An account key (ARM_ACCESS_KEY) grants full control of the entire storage account and cannot be scoped or easily audited per user.

What is .terraform/terraform.tfstate? Is that my state?

No. It is a small metadata file init writes, recording which backend type and configuration it resolved (container, key, account). It is how later commands know which blob to read, and why a freshly cloned directory needs init first. It lives in .terraform/ and is never committed. Your actual state is in the backend.

In an interview Junior

What is a Terraform backend, and how do you move existing local state into a remote one?

The backend decides where state is stored and whether it can be locked. With no backend block it is local (a file next to the code). For a team it is a shared, lockable store - with the azurerm backend, a blob in a storage account, locked with a blob lease.

To move local state:

  1. The storage account and container must already exist (a bootstrap configuration or a script creates them).
  2. Add the backend "azurerm" { ... } block. No variables are allowed there - per-environment values go in a file passed with terraform init -backend-config=backends/dev.hcl.
  3. Run terraform init. It detects the change and asks Do you want to copy existing state to the new backend? - answer yes, or you start with an empty state next to live infrastructure.
  4. Check with terraform state list, then delete the old local terraform.tfstate (it still holds secrets).

-migrate-state moves state; -reconfigure just points at a different state without copying.

Also asked: Why can a backend block not use variables? · What is partial backend configuration? · What is the difference between init -migrate-state and init -reconfigure?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.