The problem: one design, three copies
Real teams never run just one copy of their infrastructure. They run environments: separate, complete copies of the same system for different jobs.
- dev (development): small and cheap, where engineers try changes first.
- test (or staging): closer to the real thing, where changes are checked.
- prod (production): the real one that customers use.
The three must be built from the same Terraform code - otherwise "it worked in dev" means nothing - but with different sizes, names and settings. And a mistake in dev must never be able to touch prod.
So every team has to answer one question: how do we lay out Terraform so the same code builds several environments safely? There are two families of answers. This lesson explains both, and why most teams pick the first one.
What you need to know already: root modules and module calls (13.34), terraform.tfvars and variable values (12.5), backends and the state key (13.4), partial backend configuration (13.6), module version pins with ?ref= (13.38).
directory per environment CLI workspaces
envs/dev/main.tf -> module one directory
envs/prod/main.tf -> module terraform workspace select dev|prod
each with its own backend key one backend, state per workspace
and its own tfvars values chosen with terraform.workspace
Read the two columns as two layouts:
- Directory per environment: one folder per environment. Each folder is a small root module with its own state key and its own values.
- CLI workspaces: one folder, and Terraform keeps several separate states for it. A command switches which state you are working on.
Words you will meet in this lesson
- Blast radius: how much can break if one change goes wrong. A change that can only touch dev has a small blast radius.
- Pull request (PR): a proposed change in git that teammates review before it is merged into the main branch.
- Pipeline (CI pipeline): a script that a server runs automatically on every push or PR - check the code, run
terraform plan, maybe apply. Lesson 14.15 builds one; for now, "the robot that runs Terraform for the team". - Identity / credentials: the login the pipeline (or you) uses to talk to Azure. Different identities can be given different permissions (12.3).
- Promotion: moving a change from dev to test to prod, one environment at a time.
Layout 1: a directory per environment
infra/
modules/
aks/ the reusable module (lesson 14.2 builds it)
main.tf
variables.tf
outputs.tf
versions.tf
envs/
dev/
main.tf provider + module "aks" { source = "../../modules/aks" ... }
backend.tf key = "aks/dev.tfstate"
terraform.tfvars node_count = 1, vm_size = "Standard_B2s", ...
prod/
main.tf the same module call
backend.tf key = "aks/prod.tfstate" (ideally another storage account)
terraform.tfvars node_count = 3, vm_size = "Standard_D4s_v5", zones = [...]
How to read that tree:
modules/aks/holds everything every environment shares. It is written once.envs/dev/andenvs/prod/are root modules - the folders you runterraformin. Each one is short: a provider, a backend, one module call.backend.tfgives each environment its own state key, so dev and prod have separate state files (13.4).terraform.tfvarsholds the values that differ - here, how many machines (node_count) and how big (vm_size). Terraform loads it automatically (12.5).
The module in this chapter builds a small platform for running containers; the names node_count and vm_size mean "how many virtual machines" and "which machine size". Lesson 14.2 explains the module itself.
A root is mostly a module call with values:
# envs/prod/main.tf
module "aks" {
source = "../../modules/aks"
env = "prod"
location = "westeurope"
vnet_cidr = "10.110.0.0/16"
node_count = var.node_count
vm_size = var.vm_size
sku_tier = "Standard"
}
source points at the shared module. The other lines are its inputs: the environment's name, the Azure region (westeurope, a datacentre area), the network address range (a /16 - CIDR from 8.3), and sizing that comes from terraform.tfvars through variables.
Running it is the normal workflow, from inside the environment's folder:
cd envs/prod
terraform init
terraform plan -out=tfplan
cd envs/prod picks the environment. terraform init connects to prod's backend and downloads providers. terraform plan -out=tfplan computes the changes and saves them in a file called tfplan, so exactly that plan can be applied later (12.24).
What this layout gives you
- The environment is visible in the path.
cd envs/prodis impossible to miss. A pipeline can have one job per folder. A PR that only changes files underenvs/prod/is obviously a prod change. - Separate backends, separate credentials. prod's state can live in a different storage account (Azure's file storage service, 13.4) that the dev pipeline's identity cannot even read.
- Environments can differ on purpose. prod can have an extra resource dev does not need. dev can try a newer module version first (
?ref=v2.1.0in dev only). - Promotion is a diff. Moving a change from dev to prod is a PR that changes
envs/prod. Reviewers see it.
What it costs
Some repetition: each root repeats the provider block, the backend block and the module call. Keep the roots thin - provider, backend, one or two module calls, values - and the repetition is a few dozen lines that are meant to be read side by side.
Layout 2: CLI workspaces
A CLI workspace is a named, separate state for the same folder of code. Every folder starts with one workspace called default. The terraform workspace subcommands manage them:
terraform workspace new dev
terraform workspace new prod
terraform workspace list
default
dev
* prod
terraform workspace select dev
terraform workspace show
dev
| command | what it does |
|---|---|
terraform workspace new dev | create a workspace named dev and switch to it |
terraform workspace list | list them; * marks the one you are in |
terraform workspace select dev | switch to an existing workspace |
terraform workspace show | print the current workspace's name |
terraform workspace delete prod | delete one (rules below) |
Each workspace is a separate state for the same configuration. Where those states are stored depends on the backend:
local backend terraform.tfstate.d/<workspace>/terraform.tfstate (default: terraform.tfstate)
azurerm backend <key>env:<workspace> (default: <key>)
- With the local backend (state on your disk),
defaultuses the usualterraform.tfstate. Every other workspace gets a folder underterraform.tfstate.d/. - With the azurerm backend (state in a storage account), the workspace name is added to the key:
aks.tfstatebecomesaks.tfstateenv:prod.
The code finds out which workspace it runs in through the built-in value terraform.workspace (a string, e.g. "prod"), and picks values from a map:
locals {
sizing = {
dev = { node_count = 1, vm_size = "Standard_B2s" }
prod = { node_count = 3, vm_size = "Standard_D4s_v5" }
}
env = terraform.workspace
}
resource "azurerm_kubernetes_cluster" "main" {
name = "aks-orders-${local.env}"
default_node_pool {
node_count = local.sizing[local.env].node_count
vm_size = local.sizing[local.env].vm_size
}
}
local.sizing[local.env] looks up the current workspace's entry in the map (maps and lookups from 12.11). In workspace prod, the resource is named aks-orders-prod and gets 3 machines. (The resource type is the container platform from lesson 14.2; here it only matters that its name and size come from the workspace.)
The workspace facts the exam asks about
defaultalways exists and cannot be deleted.workspace newcreates and switches;workspace selectonly switches;workspace select -or-createdoes both (switch, creating it if missing).- You cannot delete the workspace you are in. You cannot delete one that still tracks resources, unless you add
-force- which leaves those resources running in Azure with no state tracking them ("orphans"). TF_WORKSPACE=prod(an environment variable) selects a workspace without runningselect. Handy in a pipeline, dangerous on a laptop, because nothing on screen shows it.terraform.workspacecan be used anywhere in the configuration.
Why directory-per-environment wins for environments
The argument, in the order an interviewer wants to hear it:
- Hidden state. Which environment you are about to change is recorded in a small file,
.terraform/environment, not in the code or the path. Aterraform applyin the wrong workspace looks exactly like one in the right workspace - until you read the resource names in the plan. - One backend, one set of credentials. All workspaces share the same backend block, so the same storage and the same identity. You cannot give prod stronger protection than dev.
- Same code, forced. Every environment runs the same code at the same version. You cannot try a new module version in dev only. Giving prod a resource dev does not have needs
count = terraform.workspace == "prod" ? 1 : 0sprinkled through the code. - Review and pipelines. A PR cannot show "this touches prod only". A pipeline cannot give prod its own approval step without extra logic.
- HashiCorp's own documentation says CLI workspaces are not a suitable way to separate environments that need separate credentials and access controls.
Where workspaces are fine: short-lived copies of the same environment - a preview per feature branch, a copy for a load test, a sandbox per developer - where sharing the backend and the code is the whole point.
Two things called "workspace"
HCP Terraform (HashiCorp's hosted service for running Terraform, lesson 14.28) also has workspaces, and they are a different thing. An HCP workspace is closer to a directory with its own state, variables, permissions and run history. Teams on HCP Terraform usually create one HCP workspace per environment per configuration - which is directory-per-environment under another name. The exam likes to test that the two "workspaces" are not the same.
Keeping environments from drifting apart
"Drift apart" here means the environments slowly stop matching - someone adds a setting to prod's root and forgets dev. Directory-per-environment trades a bit of repetition for clarity; manage that risk like this:
- Thin roots, fat modules. Everything that must be identical lives in the module. The root holds only what differs.
- Values in tfvars, not in code. One
terraform.tfvarsper environment, reviewed side by side. - Same module version everywhere by default. A different version per environment only while a change is being promoted.
- A drift job per environment. A scheduled
terraform plan -detailed-exitcodeevery night (exit 2 = something differs; 12.26) so a hand-made change in one environment gets noticed. You build it in 14.19.
Backends per environment
Each environment points at its own state:
# envs/dev/backend.tf
terraform {
backend "azurerm" {
resource_group_name = "rg-tfstate"
storage_account_name = "sttfstatesysop"
container_name = "tfstate"
key = "aks/dev.tfstate"
}
}
The four settings say where the state file lives: the resource group (Azure's folder for resources), the storage account, the container inside it, and the file name (key). Only key differs between environments here.
Two ways to write it: a literal backend.tf per environment (simple, visible), or an empty backend "azurerm" {} plus terraform init -backend-config=backends/<env>.hcl in the pipeline (partial configuration, 13.6). The rule is the same either way: one key per environment, never shared - and for production, ideally a separate storage account with its own permissions.
Promotion: the same change, one environment at a time
With a directory per environment, promoting a module change is an explicit edit in each environment - which is the point:
# envs/dev/main.tf - try it here first
module "aks" {
source = "git::https://github.com/acme/tf-modules.git//aks?ref=v2.3.0"
}
# envs/prod/main.tf - still on the version that is known to work
module "aks" {
source = "git::https://github.com/acme/tf-modules.git//aks?ref=v2.2.1"
}
Both call the same module from a git repository (13.38): //aks is the folder inside the repo, ?ref=v2.3.0 the git tag (a named version). dev is on the new version, prod still on the old one. The rollout is three PRs:
1. PR: bump dev to v2.3.0 plan dev, apply dev, watch it for a day
2. PR: bump test to v2.3.0 same
3. PR: bump prod to v2.3.0 plan shows exactly what dev and test already did
With local-path modules (source = "../../modules/aks"), every environment picks up a module change at the same time. That is fine early on, but then the environments no longer protect each other. Moving shared modules to their own repository with tags is the usual next step.
Keeping the roots thin
Everything that differs goes in tfvars; everything else is the module:
envs/prod/
backend.tf key = "aks/prod.tfstate" (or partial config from the pipeline)
main.tf provider + one module call, ~30 lines
terraform.tfvars the sizing
.terraform.lock.hcl
(.terraform.lock.hcl is the provider lock file from 12.3 - one per root.)
A reviewer compares two environments with diff, which prints the lines that differ between two files:
diff envs/dev/terraform.tfvars envs/prod/terraform.tfvars
If the diff of the two main.tf files shows more than the module ref, something environment-specific leaked into the root and belongs in the module or in tfvars.
Terragrunt, in one paragraph
Terragrunt is a separate tool that wraps Terraform. It generates the repeated parts of each root (backend block, provider block, inputs) from a tree of terragrunt.hcl files, and can run one command across many roots. It solves exactly the repetition problem of directory-per-environment. Many teams use it; many others find thin roots plus a shared pipeline script enough. It is not on the Associate exam.
What you can now do:
- Lay out a repository with
modules/and one thin root per environment. - Use
terraform workspace new/select/list/show/delete, and say where each workspace's state lives. - Argue for directory-per-environment in five points, and name where workspaces do fit.