OnCallReady

Chapter 13 Terraform: State & Modules

What state is and where it lives, locking, drift, import, and turning a flat configuration into modules without destroying anything.

In plain words

Imagine a library where the books on the shelves are the real cloud resources, and the librarian keeps a card catalogue saying "this card is that book". If the catalogue is lost, the librarian cannot tell which books belong to the library and which were dropped off by strangers. If two librarians write in the catalogue at the same time, cards get overwritten. And when the library grows, it copies the same shelf layout into every branch using a standard plan.

Terraform state is the card catalogue: it maps azurerm_subnet.this["app"] to a real Azure ID. A backend is the locked cabinet where the catalogue lives, locking stops two librarians writing at once, import adds a card for a book already on the shelf, and modules are the standard shelf plans reused in every branch.

Why it matters on call

State problems are what turn a routine Terraform change into an incident: an empty remote state next to live infrastructure after a botched init, a stale lock blocking every pipeline, a refactor that plans to destroy the production VNet because an address changed, a secret found in plaintext in a state blob someone shared. On a platform team you will also be the one designing the module that ten product teams call, and the one splitting a monolithic state without downtime.

This chapter sits right after the language basics because everything here is about identity and ownership, not syntax. It covers the state, modules and import objectives of the Terraform Associate exam, and interviewers use "what is in state and how do you protect it?" and "how do you refactor without destroying?" to separate people who have run Terraform on a team from those who have followed tutorials.

Lessons

  1. State: what it is, why it holds secrets, where it must live
  2. Backends: where state lives, and how init moves it
  3. Locking: what it prevents, and how to handle a held lock
  4. The state commands: list, show, mv, rm, pull, push, replace-provider
  5. Drift, import, and the commands that replaced taint and refresh
  6. Import: adopting what already exists
  7. Across states: sharing values, and moving objects between states
  8. Modules: layout, inputs, sources and refactoring into them
  9. Designing modules other teams can use
  10. Refactoring safely: moved, removed and lifecycle
  11. Testing modules with terraform test

43 hands-on labs (missions, incidents and drills) run in the terminal: Open this chapter in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.

Questions people ask

Why can Terraform not just look at Azure instead of keeping state?

Because Azure cannot say which object belongs to which block. Two storage accounts with similar names, one created by Terraform and one by hand, look the same from the API. State records the mapping, the dependencies needed to destroy things in the right order after their blocks are gone, cached attributes for performance, and one shared record for the team. Tools without a state file, like Bicep, rely on the platform tracking deployments instead.

Is a module the same as a function in a programming language?

Close enough to be useful. Inputs are variables, return values are outputs, and the inside is sealed: a module sees only its own variables, locals and resources. The difference is that a module's resources have addresses in state, so renaming things inside it or changing how it is called changes identity, and that can destroy real objects unless you ship moved blocks.

How many state files should a team have?

At least one per environment, so dev cannot break prod, and more when lifecycles or ownership differ: the hub network changes monthly and belongs to the platform team, apps change daily and belong to product teams. A state is a blast radius, a lock queue and a permission boundary. Signs it is too big: slow plans, teams waiting on each other's locks, unrelated changes showing up in your plan.

Do I need HCP Terraform to work on a team?

No. A remote backend such as an Azure Storage blob with RBAC, versioning and lease-based locking, plus a pipeline that plans and applies, covers the essentials. HCP Terraform adds remote runs, run history, policy checks, a private module registry, variable sets and drift detection as a product. Many teams use plain backends and CI; the exam covers both.

What is the most dangerous state operation?

Anything that rewrites state without a plan: state push -force over a newer state, -lock=false on shared state, force-unlock while the holder is still running, state rm while the block is still in the configuration. Each can create orphans or duplicates that only surface on a later apply. The safer forms are declarative and reviewable: moved, removed and import blocks.