OnCallReady

Lesson 14.1 · Terraform in Real Life & the Associate Exam · 27 min read

Environments: directory per environment vs workspaces

In plain words

Imagine a family with three kids who each have their own bedroom with their own door and their own key. The rooms have the same furniture from the same catalogue, but the teenager's bed is bigger. You always know whose room you are in because the name is on the door. The alternative is one room with a sign on the wall you flip to say whose room it is today; forget to flip it, and you tidy the wrong kid's things.

Directory per environment is the separate bedrooms: envs/dev and envs/prod, each with its own backend key and tfvars, calling the same module. CLI workspaces are the flip-sign: one directory, terraform workspace select prod, and terraform.workspace in the code. For environments, separate bedrooms win.

The problem: one design, three copies

Real teams never run just one copy of their infrastructure. They run environments: separate, complete copies of the same system for different jobs.

The three must be built from the same Terraform code - otherwise "it worked in dev" means nothing - but with different sizes, names and settings. And a mistake in dev must never be able to touch prod.

So every team has to answer one question: how do we lay out Terraform so the same code builds several environments safely? There are two families of answers. This lesson explains both, and why most teams pick the first one.

What you need to know already: root modules and module calls (13.34), terraform.tfvars and variable values (12.5), backends and the state key (13.4), partial backend configuration (13.6), module version pins with ?ref= (13.38).

directory per environment          CLI workspaces
envs/dev/main.tf  -> module         one directory
envs/prod/main.tf -> module         terraform workspace select dev|prod
each with its own backend key       one backend, state per workspace
and its own tfvars                  values chosen with terraform.workspace

Read the two columns as two layouts:

Words you will meet in this lesson

Layout 1: a directory per environment

infra/
  modules/
    aks/                 the reusable module (lesson 14.2 builds it)
      main.tf
      variables.tf
      outputs.tf
      versions.tf
  envs/
    dev/
      main.tf            provider + module "aks" { source = "../../modules/aks" ... }
      backend.tf         key = "aks/dev.tfstate"
      terraform.tfvars   node_count = 1, vm_size = "Standard_B2s", ...
    prod/
      main.tf            the same module call
      backend.tf         key = "aks/prod.tfstate"   (ideally another storage account)
      terraform.tfvars   node_count = 3, vm_size = "Standard_D4s_v5", zones = [...]

How to read that tree:

The module in this chapter builds a small platform for running containers; the names node_count and vm_size mean "how many virtual machines" and "which machine size". Lesson 14.2 explains the module itself.

A root is mostly a module call with values:

# envs/prod/main.tf
module "aks" {
  source = "../../modules/aks"

  env          = "prod"
  location     = "westeurope"
  vnet_cidr    = "10.110.0.0/16"
  node_count   = var.node_count
  vm_size      = var.vm_size
  sku_tier     = "Standard"
}

source points at the shared module. The other lines are its inputs: the environment's name, the Azure region (westeurope, a datacentre area), the network address range (a /16 - CIDR from 8.3), and sizing that comes from terraform.tfvars through variables.

Running it is the normal workflow, from inside the environment's folder:

cd envs/prod
terraform init
terraform plan -out=tfplan

cd envs/prod picks the environment. terraform init connects to prod's backend and downloads providers. terraform plan -out=tfplan computes the changes and saves them in a file called tfplan, so exactly that plan can be applied later (12.24).

What this layout gives you

What it costs

Some repetition: each root repeats the provider block, the backend block and the module call. Keep the roots thin - provider, backend, one or two module calls, values - and the repetition is a few dozen lines that are meant to be read side by side.

Layout 2: CLI workspaces

A CLI workspace is a named, separate state for the same folder of code. Every folder starts with one workspace called default. The terraform workspace subcommands manage them:

terraform workspace new dev
terraform workspace new prod
terraform workspace list
  default
  dev
* prod
terraform workspace select dev
terraform workspace show
dev
commandwhat it does
terraform workspace new devcreate a workspace named dev and switch to it
terraform workspace listlist them; * marks the one you are in
terraform workspace select devswitch to an existing workspace
terraform workspace showprint the current workspace's name
terraform workspace delete proddelete one (rules below)

Each workspace is a separate state for the same configuration. Where those states are stored depends on the backend:

local backend      terraform.tfstate.d/<workspace>/terraform.tfstate  (default: terraform.tfstate)
azurerm backend    <key>env:<workspace>                                (default: <key>)

The code finds out which workspace it runs in through the built-in value terraform.workspace (a string, e.g. "prod"), and picks values from a map:

locals {
  sizing = {
    dev  = { node_count = 1, vm_size = "Standard_B2s" }
    prod = { node_count = 3, vm_size = "Standard_D4s_v5" }
  }
  env = terraform.workspace
}

resource "azurerm_kubernetes_cluster" "main" {
  name = "aks-orders-${local.env}"
  default_node_pool {
    node_count = local.sizing[local.env].node_count
    vm_size    = local.sizing[local.env].vm_size
  }
}

local.sizing[local.env] looks up the current workspace's entry in the map (maps and lookups from 12.11). In workspace prod, the resource is named aks-orders-prod and gets 3 machines. (The resource type is the container platform from lesson 14.2; here it only matters that its name and size come from the workspace.)

The workspace facts the exam asks about

Why directory-per-environment wins for environments

The argument, in the order an interviewer wants to hear it:

  1. Hidden state. Which environment you are about to change is recorded in a small file, .terraform/environment, not in the code or the path. A terraform apply in the wrong workspace looks exactly like one in the right workspace - until you read the resource names in the plan.
  2. One backend, one set of credentials. All workspaces share the same backend block, so the same storage and the same identity. You cannot give prod stronger protection than dev.
  3. Same code, forced. Every environment runs the same code at the same version. You cannot try a new module version in dev only. Giving prod a resource dev does not have needs count = terraform.workspace == "prod" ? 1 : 0 sprinkled through the code.
  4. Review and pipelines. A PR cannot show "this touches prod only". A pipeline cannot give prod its own approval step without extra logic.
  5. HashiCorp's own documentation says CLI workspaces are not a suitable way to separate environments that need separate credentials and access controls.

Where workspaces are fine: short-lived copies of the same environment - a preview per feature branch, a copy for a load test, a sandbox per developer - where sharing the backend and the code is the whole point.

Two things called "workspace"

HCP Terraform (HashiCorp's hosted service for running Terraform, lesson 14.28) also has workspaces, and they are a different thing. An HCP workspace is closer to a directory with its own state, variables, permissions and run history. Teams on HCP Terraform usually create one HCP workspace per environment per configuration - which is directory-per-environment under another name. The exam likes to test that the two "workspaces" are not the same.

Keeping environments from drifting apart

"Drift apart" here means the environments slowly stop matching - someone adds a setting to prod's root and forgets dev. Directory-per-environment trades a bit of repetition for clarity; manage that risk like this:

Backends per environment

Each environment points at its own state:

# envs/dev/backend.tf
terraform {
  backend "azurerm" {
    resource_group_name  = "rg-tfstate"
    storage_account_name = "sttfstatesysop"
    container_name       = "tfstate"
    key                  = "aks/dev.tfstate"
  }
}

The four settings say where the state file lives: the resource group (Azure's folder for resources), the storage account, the container inside it, and the file name (key). Only key differs between environments here.

Two ways to write it: a literal backend.tf per environment (simple, visible), or an empty backend "azurerm" {} plus terraform init -backend-config=backends/<env>.hcl in the pipeline (partial configuration, 13.6). The rule is the same either way: one key per environment, never shared - and for production, ideally a separate storage account with its own permissions.

Promotion: the same change, one environment at a time

With a directory per environment, promoting a module change is an explicit edit in each environment - which is the point:

# envs/dev/main.tf - try it here first
module "aks" {
  source = "git::https://github.com/acme/tf-modules.git//aks?ref=v2.3.0"
}

# envs/prod/main.tf - still on the version that is known to work
module "aks" {
  source = "git::https://github.com/acme/tf-modules.git//aks?ref=v2.2.1"
}

Both call the same module from a git repository (13.38): //aks is the folder inside the repo, ?ref=v2.3.0 the git tag (a named version). dev is on the new version, prod still on the old one. The rollout is three PRs:

1. PR: bump dev to v2.3.0      plan dev, apply dev, watch it for a day
2. PR: bump test to v2.3.0     same
3. PR: bump prod to v2.3.0     plan shows exactly what dev and test already did

With local-path modules (source = "../../modules/aks"), every environment picks up a module change at the same time. That is fine early on, but then the environments no longer protect each other. Moving shared modules to their own repository with tags is the usual next step.

Keeping the roots thin

Everything that differs goes in tfvars; everything else is the module:

envs/prod/
  backend.tf         key = "aks/prod.tfstate" (or partial config from the pipeline)
  main.tf            provider + one module call, ~30 lines
  terraform.tfvars   the sizing
  .terraform.lock.hcl

(.terraform.lock.hcl is the provider lock file from 12.3 - one per root.)

A reviewer compares two environments with diff, which prints the lines that differ between two files:

diff envs/dev/terraform.tfvars envs/prod/terraform.tfvars

If the diff of the two main.tf files shows more than the module ref, something environment-specific leaked into the root and belongs in the module or in tfvars.

Terragrunt, in one paragraph

Terragrunt is a separate tool that wraps Terraform. It generates the repeated parts of each root (backend block, provider block, inputs) from a tree of terragrunt.hcl files, and can run one command across many roots. It solves exactly the repetition problem of directory-per-environment. Many teams use it; many others find thin roots plus a shared pipeline script enough. It is not on the Associate exam.

What you can now do:

Why it helps

You will join a team that has picked one of these, and you need to know the consequences. Situations: someone applies to prod because their shell was still in the prod workspace, and the plan looked identical to dev's; you can explain why directory per environment makes that mistake visible. A security review asks why the dev pipeline identity can read prod state; with workspaces sharing a backend, it can. You want to test a new module version in dev only; with directories, that is one ?ref= change. This is also a classic interview debate, and "workspace" meaning two different things in CLI and HCP Terraform is an exam trap.

FAQ

Are CLI workspaces bad?

No, they are a tool for a specific job: several copies of the same configuration that share a backend and credentials, like feature-branch previews, load-test copies or developer sandboxes. They are unsuitable for environments that need separate access control, different code versions or visible separation, which HashiCorp's documentation itself says. For dev, test and prod, use separate directories.

Where do workspace states live?

Each workspace is a separate state for the same configuration. With the local backend, non-default workspaces live under terraform.tfstate.d/<workspace>/terraform.tfstate. With azurerm, the blob is <key>env:<workspace>, while default uses <key> itself. The currently selected workspace is recorded in .terraform/environment, which is exactly the hidden state that makes wrong-workspace applies possible.

Is an HCP Terraform workspace the same as a CLI workspace?

No. An HCP workspace is closer to a directory: its own state, variables, permissions, run history and settings. Teams usually create one HCP workspace per configuration per environment, which is directory per environment by another name. The only overlap: a cloud block with workspaces { tags = [...] } lets terraform workspace select switch between HCP workspaces.

Doesn't directory per environment mean duplicated code?

Some, and deliberately little. Keep roots thin: provider, backend, one or two module calls and values, about 30 lines. Everything that must be identical lives in the module. The duplication is a few dozen lines meant to be read side by side; if diff envs/dev/main.tf envs/prod/main.tf shows more than the module ref, something leaked into the root. Terragrunt can generate the boilerplate if it grows.

How do I promote a change from dev to prod?

With shared modules in their own repo and tags, promotion is one PR per environment bumping the module ref: dev to v2.3.0, apply, watch; then test; then prod, whose plan should show exactly what the earlier environments already did. With local-path modules, every environment picks up changes at once, which is simpler early on but means environments no longer protect each other.

In an interview Mid

How do you manage dev and prod with Terraform: separate directories or workspaces?

Directory per environment for real environments: shared code in modules/, one thin root per environment (envs/dev, envs/prod) with its own backend key and its own terraform.tfvars.

Why not CLI workspaces (terraform workspace new prod, one folder, one state per workspace):

  1. Hidden state - which environment you are in lives in .terraform/environment, not in the path. An apply in the wrong workspace looks like the right one.
  2. One backend, one set of credentials for all workspaces - prod cannot get stronger protection than dev.
  3. Same code, forced - you cannot try a new module version in dev only; differences become terraform.workspace == "prod" ? ... : ... everywhere.
  4. A PR or a pipeline cannot show or gate "this touches prod".

Workspaces are fine for short-lived copies of one environment: a preview per branch, a load-test copy. And promotion becomes visible: bump the module ?ref= in dev, then test, then prod - three PRs.

Also asked: What are Terraform workspaces, and when would you use them? · How do you keep several environments from drifting apart? · What is the difference between a CLI workspace and an HCP Terraform workspace?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.