OnCallReady

Terraform: State & Modules: interview questions

The question you are most likely to get for each topic, a model answer, and what else comes up. From chapter 13 of the course.

What is Terraform state, and how do you manage it safely in a team? Mid

State is the JSON file that maps every resource address in the code to the real object it manages (azurerm_resource_group.main -> its resource ID), plus metadata like dependencies. It also holds every attribute in plain text, secrets included.

Managing it safely:

Also asked: What is a Terraform backend? · What is configuration drift, and how do you detect it? · What is a Terraform module, and why would you use one?

Why does Terraform need a state file, and why is it sensitive? Junior

The configuration says what should exist and the cloud says what does exist, but neither says which real object belongs to which block. State is that mapping: each resource address to its resource ID. It also records dependencies (so a deleted block is still destroyed in the right order), caches attributes for performance, and tracks a serial (write counter) and lineage (identity) that protect against overwriting newer state with older.

It is sensitive because it holds every attribute the provider returned, in plain text - generated passwords, access keys, connection strings. sensitive = true only hides values on screen, not in state.

So: never commit it (.gitignore covers *.tfstate), keep it in a remote backend that is encrypted, versioned and access-controlled, and never edit it by hand. terraform state pull | jq is how you read it.

Also asked: What happens if two people run terraform apply at the same time without locking? · What is the difference between serial and lineage in a state file? · Why should you have one state per environment?

Learn it: 13.1 State: what it is, why it holds secrets, where it must live

What is a Terraform backend, and how do you move existing local state into a remote one? Junior

The backend decides where state is stored and whether it can be locked. With no backend block it is local (a file next to the code). For a team it is a shared, lockable store - with the azurerm backend, a blob in a storage account, locked with a blob lease.

To move local state:

  1. The storage account and container must already exist (a bootstrap configuration or a script creates them).
  2. Add the backend "azurerm" { ... } block. No variables are allowed there - per-environment values go in a file passed with terraform init -backend-config=backends/dev.hcl.
  3. Run terraform init. It detects the change and asks Do you want to copy existing state to the new backend? - answer yes, or you start with an empty state next to live infrastructure.
  4. Check with terraform state list, then delete the old local terraform.tfstate (it still holds secrets).

-migrate-state moves state; -reconfigure just points at a different state without copying.

Also asked: Why can a backend block not use variables? · What is partial backend configuration? · What is the difference between init -migrate-state and init -reconfigure?

Learn it: 13.4 Backends: where state lives, and how init moves it

A Terraform pipeline fails with "Error acquiring the state lock". What do you do? Junior

That error is locking working: another run holds the lock on this state. Never just use -lock=false.

  1. Read the Lock Info: ID, Path (which state), Who (a pipeline agent or a laptop), Operation, Created.
  2. Find the holder - that pipeline run or that colleague.
  3. If it is still running, wait. In pipelines, -lock-timeout=5m makes the second run wait instead of failing, and the CI setting for one run per state at a time avoids the queue.
  4. If the holder is confirmed dead (cancelled run, crashed agent): terraform force-unlock <ID> with exactly that ID.
  5. Assume the dead run did part of its work: an object it created before dying may be an orphan (it exists but is not in state). Import it; do not delete it.

A lock does not protect against two states managing the same object, changes made by hand, or a stale saved plan.

Also asked: What is state locking and why does it matter? · Which Terraform commands take the state lock? · What can go wrong if you force-unlock a lock whose run is still alive?

Learn it: 13.11 Locking: what it prevents, and how to handle a held lock

You renamed a resource in code and the plan wants to destroy and recreate it. How do you fix it? Mid

Terraform matches config to state by address. A rename looks like one object deleted and another added, so it plans a destroy and a create of the same thing.

Tell it about the rename:

moved {
  from = azurerm_storage_account.logs
  to   = azurerm_storage_account.diagnostics
}

Either way the real object keeps its ID; only its address changes. The check is the next plan: has moved to, and 0 to add, 0 to destroy.

Also asked: What do terraform state list and terraform state show tell you? · What does terraform state rm do, and what does it leave behind? · When would you use terraform state push?

Learn it: 13.15 The state commands: list, show, mv, rm, pull, push, replace-provider

What is configuration drift in Terraform, and how do you detect and handle it? Junior

Drift is any difference between state and reality: someone changed a firewall rule by hand during an incident, a policy added tags, an autoscaler changed a count.

Detect: every plan refreshes first, and drift appears as Objects have changed outside of Terraform above the plan. terraform plan -refresh-only shows only the drift. A nightly plan -detailed-exitcode (exit 2 = changes) finds it without waiting for the next deploy.

Handle, per object:

Never let a pipeline auto-approve a plan that opens with that drift note: half of drift is a deliberate fix the code has not caught up with.

Also asked: What is the difference between terraform plan -refresh-only and a normal plan? · What replaced terraform taint, and why? · What does ignore_changes do, and when is it the wrong answer?

Learn it: 13.21 Drift, import, and the commands that replaced taint and refresh

How do you bring an existing, hand-built resource under Terraform management? Junior

Import it: write a state entry that says "this address manages that existing object". Nothing in the cloud changes.

  1. Write the resource block (or let terraform plan -generate-config-out=generated.tf write a starting one).
  2. Find the object's resource ID (the provider docs' Import section shows the format).
  3. Add an import block - reviewed in a PR, part of plan and apply:
import {
  to = azurerm_storage_account.reports
  id = "/subscriptions/.../storageAccounts/streports"
}

(The older terraform import ADDRESS ID writes state immediately with no plan.)

  1. Plan until it says 1 to import, 0 to change. Any ~ or -/+ means your block does not match reality, and the apply would change - or replace - what you just adopted.

Import takes ownership; a data source only reads something someone else owns.

Also asked: What is the difference between import and a data source? · What does -generate-config-out do, and what do you still have to do afterwards? · What is an orphaned resource, and how does one appear?

Learn it: 13.25 Import: adopting what already exists

How do you share values between Terraform configurations that have separate states? Mid

Three options:

Moving an object between states is different: a removed block with destroy = false in the source, then an import block in the destination, both in one change window, and the destination plan must say 1 to import, 0 to change.

Also asked: Another team needs the ID of a subnet your Terraform manages. How do you give it to them? · How would you split a large Terraform state into smaller ones? · Why should two states never manage the same object?

Learn it: 13.30 Across states: sharing values, and moving objects between states

What is a Terraform module and why would you use one? Junior

A module is a directory of .tf files used as a reusable unit with inputs and outputs - like a function. The directory you run Terraform in is the root module; it calls child modules with a module block:

module "network" {
  source        = "./modules/network"
  address_space = ["10.40.0.0/16"]
}

Why: write it once and every team gets the same, consistent, reviewed resources; a fix is made in one place. Pin versions, because module versions are not in the lock file.

Also asked: What module sources can you use, and how do you version them? · Why is depends_on on a module usually a bad idea? · How do you pass an aliased provider into a module?

Learn it: 13.34 Modules: layout, inputs, sources and refactoring into them

How do you design a Terraform module for other teams to use? Mid

Treat it as a small product with an interface:

Also asked: What makes a good module input variable? · How do you roll out a breaking change in a module used by many teams? · Why should a module not contain a provider block?

Learn it: 13.41 Designing modules other teams can use

What is a moved block, and when would you use it instead of terraform state mv? Mid

A moved block tells Terraform that the object at from now lives at to, so a refactor carries the old identity across instead of planning a destroy and a create:

moved {
  from = azurerm_subnet.app
  to   = module.network.azurerm_subnet.this["app"]
}

It covers renames, count to for_each, moving into or between modules. The plan shows has moved to and 0 to add, 0 to destroy.

Use moved by default: it is in the code, reviewed in a PR, applied by every environment's next apply, and a shared module ships it to every caller who upgrades. state mv runs from your shell against one state, unreviewed - for a one-off repair.

Related: a removed block (with destroy = false) is the reviewable state rm, and lifecycle settings - prevent_destroy, create_before_destroy, ignore_changes, replace_triggered_by - protect objects that really must change.

Also asked: What does create_before_destroy do, and why does it often fail on cloud resources with unique names? · What does prevent_destroy protect against, and what does it not? · How do you refactor a flat configuration into modules without destroying anything?

Learn it: 13.46 Refactoring safely: moved, removed and lifecycle

How do you test a Terraform module? Junior

With the native framework, terraform test (1.6+). Test files end in .tftest.hcl, in the module directory or tests/:

variables {
  subnets = { app = { cidr = "10.40.1.0/24" } }
}

run "creates_one_subnet_per_entry" {
  command = plan
  assert {
    condition     = length(azurerm_subnet.this) == 1
    error_message = "Expected one subnet per entry."
  }
}

run "rejects_bad_cidr" {
  command = plan
  variables { subnets = { app = { cidr = "10.40.1.0/33" } } }
  expect_failures = [var.subnets]
}

terraform test exits non-zero on failure, so CI runs it on every PR before a module version is tagged.

Also asked: What would you put in a test suite for a network module? · What is the difference between a plan run and an apply run in terraform test? · Where do Terraform tests fit in a CI job?

Learn it: 13.53 Testing modules with terraform test

Practise these answers with flashcards and labs Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.