The problem it solves
What you need to know already: files and directories (1.3, 1.5), editing a file with a heredoc or nano (6.18), git and commits (1.17), JSON (7.11), and the idea of a server you reach over the network (Ch 8-9). No cloud knowledge is assumed: this lesson explains the little you need.
Picture a team that builds its servers by clicking through a website. It works on the day. Six months later nobody remembers why one firewall rule exists, the test copy is "almost" like the real one, and rebuilding everything after a disaster is a week of guessing. Nothing was written down, so nothing can be reviewed, repeated or undone.
This chapter fixes that with Terraform, a tool that builds infrastructure from text files.
A few words first
- Infrastructure: the things your applications run on - machines, networks, disks, databases, DNS names, firewall rules.
- Cloud: a company that rents you infrastructure over the internet. You do not rack a server; you ask the provider's API (a web address that accepts requests, like the HTTP APIs from 9.21) and a minute later the thing exists. You pay while it exists.
- Azure: Microsoft's cloud. The labs in this course use it because it is what the team uses at work, but you do not need to know Azure - every Azure thing gets a one-line explanation where it first appears. (Azure itself has its own chapters much later.)
- Portal: the cloud's website, where you can click things into existence. That clicking is exactly what Terraform replaces.
Infrastructure as Code
Infrastructure as Code (IaC) means the infrastructure is described in text files that live in git. Everything you know from frontend work carries over: pull requests, review, CI, history, blame, revert.
What IaC buys you, in the words the exam and interviewers use:
reproducible the same code builds the same environment, again and again
reviewable a change is a diff someone reads before it happens
versioned git log is the change history of your infrastructure
self-documenting the code IS the inventory - no stale wiki page
automatable a script on a server can run it; no human clicking
consistent dev, test and prod come from one definition with different inputs
(dev, test, prod: copies of the same system - dev for developers to break, test for checking, prod = production, the one real users hit.)
Declarative vs imperative
An imperative tool is a list of steps: create this, then attach that. A bash script (Ch 6) is imperative. Run it twice and you get two of everything, or an error the second time.
A declarative tool is a description of the end state. You write "a group called rg-orders-dev in westeurope must exist". The tool compares that with what exists and does whatever is needed - which is nothing, if it already exists. Terraform is declarative. You have met the idea already: a systemd unit file (2.1) describes a service, and systemd makes it so.
Here is the smallest Terraform example. Read it line by line below.
resource "azurerm_resource_group" "orders" {
name = "rg-orders-dev"
location = "westeurope"
}
resourcemeans "Terraform creates and owns this thing"."azurerm_resource_group"is the kind of thing. In Azure a resource group is a named folder that every other Azure thing must live in; delete the folder and everything in it goes too."orders"is the nickname you use for it inside your files.nameis what Azure will call it;locationis the region, the city of Microsoft data centres it lives in (westeurope= the Netherlands).
That block does not say "create". It says "this must exist". Whether Terraform creates, updates, replaces or leaves it alone depends on what is already there.
The property that falls out of this is idempotence: applying the same configuration twice changes nothing the second time. It is what makes it safe to run Terraform on every commit.
Mutable vs immutable
When something must change, a tool can edit it in place (mutable) or throw it away and build a new one (immutable). Terraform does both, and which one is decided by the provider, per attribute (per setting). Tags - free-form labels like env = "dev" - on a resource group change in place. The resource group's name cannot change, so a new name means destroy and create.
You will learn to read that in every plan. Missing it is the most expensive mistake in this chapter: "destroy and create" on something full of data means the data is gone.
How Terraform is built
Terraform is two kinds of program working together:
terraform (core) one Go binary: reads your .tf files, builds a graph,
| diffs, keeps state, runs the workflow
| gRPC plugin protocol (how the two programs talk to each other)
v
providers separate binaries, one per platform:
hashicorp/azurerm Microsoft Azure
hashicorp/random random strings, passwords, pet names
hashicorp/local files on this machine
integrations/github GitHub repositories, teams
... thousands more on registry.terraform.io
A provider is a plugin that knows one platform's API. Core knows nothing about Azure. Core reads configuration, builds a dependency graph (which thing must exist before which), compares what you want with what it recorded, and calls providers. Every resource type (azurerm_subnet) and every API call belongs to a provider.
That split is why one tool can manage Azure, GitHub, DNS and many others with the same workflow - and why terraform init has to download providers before anything else works.
The third piece is state: a JSON file (7.11) that maps every resource in your code to the real object it manages, e.g. azurerm_resource_group.orders = /subscriptions/.../resourceGroups/rg-orders-dev. (That long path is the object's ID in Azure; a subscription is the Azure billing account everything is created under.) Ch 13 is all about state.
Terraform vs the other tools
This is exam objective 1 and a standard interview warm-up. You do not need to know these tools; just the column that says how each one works.
tool kind scope state
Terraform declarative, HCL any provider state file (you manage it)
OpenTofu fork of Terraform 1.5 any provider same model, MPL licensed
Bicep / ARM declarative Azure only Azure tracks deployments
CloudFormation declarative AWS only AWS tracks stacks
Pulumi declarative, real any provider state (service or backend)
languages (TS, Go, C#)
Ansible procedural tasks, config of hosts no state; re-runs tasks
mostly idempotent
(AWS is Amazon's cloud. A fork is a copy of a project that continues separately. HCL is Terraform's language - next section.)
Two distinctions to be able to make out loud:
- Provisioning vs configuration management. Provisioning = creating the machine, the network, the disk. Configuration management = setting up what runs inside the machine (packages, files, services - the Ch 1-2 work). Terraform provisions; Ansible (or a script, or a prepared image) configures. They are complements, not competitors.
- Multi-cloud means one workflow, not one configuration. Terraform does not let you write one resource that deploys to Azure or AWS. An Azure virtual machine and an AWS virtual machine are different resource types with different arguments. What you get is the same language, the same
init/plan/apply, the same state model and the same review process everywhere - and the ability to wire providers together in one configuration (create a server in Azure, then a DNS record for it at another DNS company).
Licensing, since it comes up: Terraform moved from the open-source MPL licence to the Business Source License in version 1.6 (August 2023). OpenTofu is the community fork of 1.5.x under the Linux Foundation. For using Terraform at work nothing changed; it matters to vendors building competing products.
The files
Terraform code is written in HCL (HashiCorp Configuration Language): blocks in curly braces, name = value lines, # comments. A directory of .tf files is one configuration.
Terraform reads every .tf file in the current directory (not subdirectories) and merges them. File names mean nothing to Terraform; they are for humans. The convention everyone follows:
main.tf the resources
variables.tf input variable declarations
outputs.tf output declarations
providers.tf the terraform { } block and provider configuration
versions.tf (some teams) required_version and required_providers
terraform.tfvars values for the variables - loaded automatically
locals.tf (optional) computed values
Order does not matter either: a resource can reference one defined further down or in another file. Terraform builds a graph from the references, not from the order you wrote things in.
What else appears in a working directory:
.terraform/ providers and modules downloaded by init - never commit
.terraform.lock.hcl exact provider versions + checksums - COMMIT this
terraform.tfstate local state (when there is no backend) - NEVER commit
terraform.tfstate.backup the previous state - never commit
*.tfvars values; commit the non-secret ones
(A checksum is a fingerprint of a file: if one byte changes, the checksum changes, so a tampered download is caught.)
The block types
Everything in HCL is a block. You will meet each of these in Ch 12-14; for now just recognise the names.
terraform { } settings: required versions, providers, backend
provider { } how to talk to a platform (credentials, region, features)
resource { } something Terraform CREATES and owns
data { } something it READS but does not own
variable { } an input
output { } a value to expose
locals { } computed or repeated values, internal
module { } a reusable group of resources
import { } adopt an existing resource (1.5+)
moved { } a resource was renamed - move it in state, do not destroy it (1.1+)
removed { } stop managing something without destroying it (1.7+)
check { } continuous assertions that warn, never block (1.5+)
A block has a type, zero or more labels (the quoted words after the type), and a body in braces made of arguments (name = value) and nested blocks:
resource "azurerm_storage_account" "exports" { # type, then two labels
name = "stordersexportsdev" # argument = expression
resource_group_name = azurerm_resource_group.orders.name
location = "westeurope"
account_tier = "Standard"
account_replication_type = "LRS"
blob_properties { # a nested block
versioning_enabled = true
}
}
An Azure storage account is a place to keep files in the cloud (Azure calls the files blobs). account_tier and account_replication_type = "LRS" (locally redundant: 3 copies in one data centre) are its price/safety settings. versioning_enabled keeps old versions of each file.
For a resource, the first label is the resource type (the prefix before the first underscore picks the provider: azurerm_), the second is the local name you use to refer to it. Together they form its address: azurerm_storage_account.exports.
The line resource_group_name = azurerm_resource_group.orders.name is a reference: "use the name of the resource group I called orders". That is how resources are connected, and how Terraform knows the group must exist first.
The local name exists only inside your configuration; the name argument is what Azure calls it. Renaming one does not rename the other - and renaming the local name without a moved block makes Terraform think you deleted one resource and added another.
resource vs data is the distinction to be precise about: a resource is created and destroyed by Terraform; a data source is a lookup of something that already exists. terraform destroy never deletes a data source's target.
The terraform block
The terraform { } block holds settings for Terraform itself.
terraform {
required_version = "~> 1.9"
required_providers {
azurerm = {
source = "hashicorp/azurerm"
version = "~> 4.0"
}
}
backend "azurerm" {
resource_group_name = "rg-tfstate"
storage_account_name = "sttfstateprod"
container_name = "tfstate"
key = "platform/prod.tfstate"
}
}
required_versionpins the Terraform binary.~> 1.9means "1.9 or any newer 1.x" (the next lesson explains~>). A teammate on an older version gets a clear error instead of a silently different plan.required_providerssays which providers and which versions. The next lesson is all about it.backendsays where state lives - here, in a storage account, so the whole team shares one copy. Only literal values are allowed here - no variables - because the backend is needed before variables are evaluated. Ch 13 covers backends.
In the lab, try a constraint the lab binary does not meet:
# the first mission's main.tf with required_version = ">= 1.10"
terraform plan
╷
│ Error: Unsupported Terraform Core version
│
│ on main.tf line 2, in terraform:
│ 2: required_version = ">= 1.10"
│
│ This configuration does not support Terraform version 1.9.8. To proceed,
│ either choose another supported Terraform version or update this version
│ constraint. Version constraints are normally set for good reason, so
│ updating the constraint may lead to other errors or unexpected behavior.
╵
How to read a Terraform error box: the first line is the error; on main.tf line 2 is where; the quoted line is the offending code; the paragraph is the explanation. Every Terraform error has this shape.
The workflow in one breath
These are terraform subcommands: you type terraform init, terraform plan and so on, in the directory that holds the .tf files.
write describe what you want (edit .tf files)
init download providers and modules, configure the backend
plan compare configuration with state (and reality) - show the diff
apply make the changes, record the result in state
destroy remove everything this configuration manages
plan is the part people undervalue. It is a diff against reality, and it is what makes an infrastructure change reviewable the same way a code change is. Every lesson from here on ends up in the plan output; learning to read it carefully is most of the job.
A note on this lab: the simulated box runs Terraform 1.9.8 with an azurerm provider that behaves like the real one for the resources the course uses. Nothing is really created in Azure - the "cloud" is simulated on the box (simulator). The real Associate exam (currently version 004) targets Terraform 1.12; the handful of 1.10-1.12 features that matter for it (ephemeral values, write-only arguments) are covered as concepts in Ch 14.
What you can now do:
- Say what IaC is and why declarative + idempotent makes it safe to re-run.
- Read a
resourceblock: type, local name, arguments, references. - Name the five workflow steps and the files a working directory contains.