OnCallReady

Lesson 12.1 · Terraform: Language & Workflow · 27 min read

What Terraform is for, and the blocks it is made of

In plain words

Think of a recipe card versus a photo of the finished cake. A recipe card (imperative) says "mix, then bake, then ice"; follow it twice and you have two cakes. A photo (declarative) says "the table should have this cake on it"; a smart baker looks at the table, and if the cake is already there, does nothing.

Terraform works from the photo. A resource "azurerm_resource_group" "orders" block does not mean "create"; it means "this must exist". Terraform core compares that with its state, and providers like hashicorp/azurerm make the actual Azure API calls. The block types (resource, data, variable, output, module and friends) are the vocabulary for describing that photo.

The problem it solves

What you need to know already: files and directories (1.3, 1.5), editing a file with a heredoc or nano (6.18), git and commits (1.17), JSON (7.11), and the idea of a server you reach over the network (Ch 8-9). No cloud knowledge is assumed: this lesson explains the little you need.

Picture a team that builds its servers by clicking through a website. It works on the day. Six months later nobody remembers why one firewall rule exists, the test copy is "almost" like the real one, and rebuilding everything after a disaster is a week of guessing. Nothing was written down, so nothing can be reviewed, repeated or undone.

This chapter fixes that with Terraform, a tool that builds infrastructure from text files.

A few words first

Infrastructure as Code

Infrastructure as Code (IaC) means the infrastructure is described in text files that live in git. Everything you know from frontend work carries over: pull requests, review, CI, history, blame, revert.

What IaC buys you, in the words the exam and interviewers use:

reproducible     the same code builds the same environment, again and again
reviewable       a change is a diff someone reads before it happens
versioned        git log is the change history of your infrastructure
self-documenting the code IS the inventory - no stale wiki page
automatable      a script on a server can run it; no human clicking
consistent       dev, test and prod come from one definition with different inputs

(dev, test, prod: copies of the same system - dev for developers to break, test for checking, prod = production, the one real users hit.)

Declarative vs imperative

An imperative tool is a list of steps: create this, then attach that. A bash script (Ch 6) is imperative. Run it twice and you get two of everything, or an error the second time.

A declarative tool is a description of the end state. You write "a group called rg-orders-dev in westeurope must exist". The tool compares that with what exists and does whatever is needed - which is nothing, if it already exists. Terraform is declarative. You have met the idea already: a systemd unit file (2.1) describes a service, and systemd makes it so.

Here is the smallest Terraform example. Read it line by line below.

resource "azurerm_resource_group" "orders" {
  name     = "rg-orders-dev"
  location = "westeurope"
}

That block does not say "create". It says "this must exist". Whether Terraform creates, updates, replaces or leaves it alone depends on what is already there.

The property that falls out of this is idempotence: applying the same configuration twice changes nothing the second time. It is what makes it safe to run Terraform on every commit.

Mutable vs immutable

When something must change, a tool can edit it in place (mutable) or throw it away and build a new one (immutable). Terraform does both, and which one is decided by the provider, per attribute (per setting). Tags - free-form labels like env = "dev" - on a resource group change in place. The resource group's name cannot change, so a new name means destroy and create.

You will learn to read that in every plan. Missing it is the most expensive mistake in this chapter: "destroy and create" on something full of data means the data is gone.

How Terraform is built

Terraform is two kinds of program working together:

terraform (core)         one Go binary: reads your .tf files, builds a graph,
     |                   diffs, keeps state, runs the workflow
     |  gRPC plugin protocol   (how the two programs talk to each other)
     v
providers                separate binaries, one per platform:
  hashicorp/azurerm        Microsoft Azure
  hashicorp/random         random strings, passwords, pet names
  hashicorp/local          files on this machine
  integrations/github      GitHub repositories, teams
  ...                      thousands more on registry.terraform.io

A provider is a plugin that knows one platform's API. Core knows nothing about Azure. Core reads configuration, builds a dependency graph (which thing must exist before which), compares what you want with what it recorded, and calls providers. Every resource type (azurerm_subnet) and every API call belongs to a provider.

That split is why one tool can manage Azure, GitHub, DNS and many others with the same workflow - and why terraform init has to download providers before anything else works.

The third piece is state: a JSON file (7.11) that maps every resource in your code to the real object it manages, e.g. azurerm_resource_group.orders = /subscriptions/.../resourceGroups/rg-orders-dev. (That long path is the object's ID in Azure; a subscription is the Azure billing account everything is created under.) Ch 13 is all about state.

Terraform vs the other tools

This is exam objective 1 and a standard interview warm-up. You do not need to know these tools; just the column that says how each one works.

tool                 kind                    scope            state
Terraform            declarative, HCL        any provider     state file (you manage it)
OpenTofu             fork of Terraform 1.5   any provider     same model, MPL licensed
Bicep / ARM          declarative             Azure only       Azure tracks deployments
CloudFormation       declarative             AWS only         AWS tracks stacks
Pulumi               declarative, real       any provider     state (service or backend)
                     languages (TS, Go, C#)
Ansible              procedural tasks,       config of hosts  no state; re-runs tasks
                     mostly idempotent

(AWS is Amazon's cloud. A fork is a copy of a project that continues separately. HCL is Terraform's language - next section.)

Two distinctions to be able to make out loud:

Licensing, since it comes up: Terraform moved from the open-source MPL licence to the Business Source License in version 1.6 (August 2023). OpenTofu is the community fork of 1.5.x under the Linux Foundation. For using Terraform at work nothing changed; it matters to vendors building competing products.

The files

Terraform code is written in HCL (HashiCorp Configuration Language): blocks in curly braces, name = value lines, # comments. A directory of .tf files is one configuration.

Terraform reads every .tf file in the current directory (not subdirectories) and merges them. File names mean nothing to Terraform; they are for humans. The convention everyone follows:

main.tf           the resources
variables.tf      input variable declarations
outputs.tf        output declarations
providers.tf      the terraform { } block and provider configuration
versions.tf       (some teams) required_version and required_providers
terraform.tfvars  values for the variables - loaded automatically
locals.tf         (optional) computed values

Order does not matter either: a resource can reference one defined further down or in another file. Terraform builds a graph from the references, not from the order you wrote things in.

What else appears in a working directory:

.terraform/              providers and modules downloaded by init - never commit
.terraform.lock.hcl      exact provider versions + checksums  - COMMIT this
terraform.tfstate        local state (when there is no backend) - NEVER commit
terraform.tfstate.backup the previous state                    - never commit
*.tfvars                 values; commit the non-secret ones

(A checksum is a fingerprint of a file: if one byte changes, the checksum changes, so a tampered download is caught.)

The block types

Everything in HCL is a block. You will meet each of these in Ch 12-14; for now just recognise the names.

terraform { }     settings: required versions, providers, backend
provider  { }     how to talk to a platform (credentials, region, features)
resource  { }     something Terraform CREATES and owns
data      { }     something it READS but does not own
variable  { }     an input
output    { }     a value to expose
locals    { }     computed or repeated values, internal
module    { }     a reusable group of resources
import    { }     adopt an existing resource (1.5+)
moved     { }     a resource was renamed - move it in state, do not destroy it (1.1+)
removed   { }     stop managing something without destroying it (1.7+)
check     { }     continuous assertions that warn, never block (1.5+)

A block has a type, zero or more labels (the quoted words after the type), and a body in braces made of arguments (name = value) and nested blocks:

resource "azurerm_storage_account" "exports" {   # type, then two labels
  name                     = "stordersexportsdev" # argument = expression
  resource_group_name      = azurerm_resource_group.orders.name
  location                 = "westeurope"
  account_tier             = "Standard"
  account_replication_type = "LRS"

  blob_properties {                                # a nested block
    versioning_enabled = true
  }
}

An Azure storage account is a place to keep files in the cloud (Azure calls the files blobs). account_tier and account_replication_type = "LRS" (locally redundant: 3 copies in one data centre) are its price/safety settings. versioning_enabled keeps old versions of each file.

For a resource, the first label is the resource type (the prefix before the first underscore picks the provider: azurerm_), the second is the local name you use to refer to it. Together they form its address: azurerm_storage_account.exports.

The line resource_group_name = azurerm_resource_group.orders.name is a reference: "use the name of the resource group I called orders". That is how resources are connected, and how Terraform knows the group must exist first.

The local name exists only inside your configuration; the name argument is what Azure calls it. Renaming one does not rename the other - and renaming the local name without a moved block makes Terraform think you deleted one resource and added another.

resource vs data is the distinction to be precise about: a resource is created and destroyed by Terraform; a data source is a lookup of something that already exists. terraform destroy never deletes a data source's target.

The terraform block

The terraform { } block holds settings for Terraform itself.

terraform {
  required_version = "~> 1.9"

  required_providers {
    azurerm = {
      source  = "hashicorp/azurerm"
      version = "~> 4.0"
    }
  }

  backend "azurerm" {
    resource_group_name  = "rg-tfstate"
    storage_account_name = "sttfstateprod"
    container_name       = "tfstate"
    key                  = "platform/prod.tfstate"
  }
}

In the lab, try a constraint the lab binary does not meet:

# the first mission's main.tf with required_version = ">= 1.10"
terraform plan
╷
│ Error: Unsupported Terraform Core version
│
│   on main.tf line 2, in terraform:
│    2:   required_version = ">= 1.10"
│
│ This configuration does not support Terraform version 1.9.8. To proceed,
│ either choose another supported Terraform version or update this version
│ constraint. Version constraints are normally set for good reason, so
│ updating the constraint may lead to other errors or unexpected behavior.
╵

How to read a Terraform error box: the first line is the error; on main.tf line 2 is where; the quoted line is the offending code; the paragraph is the explanation. Every Terraform error has this shape.

The workflow in one breath

These are terraform subcommands: you type terraform init, terraform plan and so on, in the directory that holds the .tf files.

write     describe what you want (edit .tf files)
init      download providers and modules, configure the backend
plan      compare configuration with state (and reality) - show the diff
apply     make the changes, record the result in state
destroy   remove everything this configuration manages

plan is the part people undervalue. It is a diff against reality, and it is what makes an infrastructure change reviewable the same way a code change is. Every lesson from here on ends up in the plan output; learning to read it carefully is most of the job.

A note on this lab: the simulated box runs Terraform 1.9.8 with an azurerm provider that behaves like the real one for the resources the course uses. Nothing is really created in Azure - the "cloud" is simulated on the box (simulator). The real Associate exam (currently version 004) targets Terraform 1.12; the handful of 1.10-1.12 features that matter for it (ephemeral values, write-only arguments) are covered as concepts in Ch 14.

What you can now do:

Why it helps

Two situations where this lesson pays off. First, a PR review: a teammate renames name = "rg-orders-dev" and the plan says must be replaced. Knowing that the provider decides in-place vs replace per attribute, you ask what lives in that resource group before anyone types yes. Second, an interview warm-up: "Terraform vs Bicep vs Ansible" and "declarative vs imperative" come up in nearly every platform interview and in exam objective 1. Knowing what core does versus what a provider does also explains everyday errors: why nothing works before terraform init, and why renaming a resource's local name without a moved block plans a destroy.

FAQ

What is the difference between the local name and the name argument?

In resource "azurerm_storage_account" "exports" { name = "stordersexportsdev" }, exports is the local name: it exists only in your code and forms the address azurerm_storage_account.exports. The name argument is what Azure calls the object. Changing the name argument usually forces a replacement in Azure. Changing the local name changes the address, so without a moved block Terraform thinks one resource was deleted and another added.

Does it matter how I name my .tf files or in what order I write blocks?

No. Terraform reads every .tf file in the current directory (not subdirectories) and merges them. main.tf, variables.tf, outputs.tf and providers.tf are conventions for humans. Order does not matter either: Terraform builds a dependency graph from references, so a resource can refer to one defined further down or in another file.

Which files do I commit and which do I ignore?

Commit the .tf files, non-secret .tfvars and .terraform.lock.hcl. Never commit .terraform/ (downloaded providers and modules, rebuilt by init), terraform.tfstate or terraform.tfstate.backup (state holds secrets in plaintext and belongs in a remote backend), or any tfvars file containing secrets. A good .gitignore for Terraform is the first thing in a new repo.

Terraform changed its licence. Should I use OpenTofu instead?

For using Terraform at work to manage your own infrastructure, the licence change (MPL to BSL in 1.6, August 2023) changed nothing. It matters to vendors building competing hosted products. OpenTofu is the Linux Foundation fork of 1.5.x and is largely compatible, with its own features since. Some organisations standardise on it; the exam and most job ads still say Terraform, and the concepts transfer almost 1:1.

Why can the backend block not use variables?

The backend is needed before Terraform evaluates anything else: it has to know where state lives to plan at all, and variables are evaluated later. So the backend block accepts only literal values. To vary it per environment you use partial configuration (terraform init -backend-config=prod.hcl), covered in Ch 13.

In an interview Junior

What is Infrastructure as Code, and why is Terraform called declarative?

Infrastructure as Code means the infrastructure - networks, machines, databases, firewall rules - is described in text files kept in git. So it is reproducible (the same code builds the same environment), reviewable (a change is a diff someone reads before it happens) and versioned (history, blame, revert).

Declarative: a resource block does not say "create"; it says "this must exist":

resource "azurerm_resource_group" "orders" {
  name     = "rg-orders-dev"
  location = "westeurope"
}

Terraform compares that with what exists and does whatever is needed - nothing, if it is already there. That makes it idempotent, safe to run on every commit. An imperative tool, like a bash script, is a list of steps: run it twice and you get two of everything or an error.

Terraform provisions (creates the infrastructure); configuration management like Ansible sets up what runs inside a machine. They complement each other.

Also asked: What is the difference between Terraform and Ansible? · What is a Terraform provider? · What does "multi-cloud" mean with Terraform, and what does it not mean?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.