OnCallReady

Lesson 14.23 · Terraform in Real Life & the Associate Exam · 22 min read

Pinning and upgrading Terraform, providers and modules

In plain words

Think of a school uniform shop. You wear the size on your label (the pin), and the shop keeps a record of exactly what you bought (the lock file). When you grow, you don't grab a random bigger size on the day of the school photo; you try the new size at home first, check the sleeves, then switch. And once you have written your name in permanent marker on the new jacket, you can't give it back.

Terraform has three things to upgrade: Terraform itself (required_version = "~> 1.9.0" plus the CI pin), providers (~> 4.14 plus .terraform.lock.hcl, moved with terraform init -upgrade), and modules (version or ?ref=). State is the permanent marker: once a newer Terraform writes it, older binaries refuse to read it. Every upgrade is a PR with a clean plan in every environment.

The problem: versions move whether you are ready or not

Terraform, the azurerm provider and the modules you use all release new versions every few weeks. Most releases are harmless. Some rename arguments, change defaults, or write state in a newer format. If versions move by accident

out in the middle of an unrelated change, usually in production.

So platform teams pin every version, and move them deliberately: one PR, a plan against every environment, dev first.

What you need to know already: version constraints (~>, >=) and the lock file .terraform.lock.hcl (12.3, 12.4), module version and ?ref= (13.38), the environments layout (14.1), terraform validate (14.6), PRs and pipelines (14.15).

Versions are written major.minor.patch (semantic versioning): 4.14.0 is major 4, minor 14, patch 0. A patch fixes bugs, a minor adds features, a major may break things - rename or remove arguments, change defaults.

Three things that get upgraded

Terraform itself     required_version, the binary in CI and on laptops
providers            required_providers constraints + .terraform.lock.hcl
modules              version / ?ref= per call

Each has its own pin and its own way of moving. The discipline is the same: move deliberately, in a PR, with a plan against every environment.

Pinning Terraform

terraform {
  required_version = "~> 1.9.0"
}
Error: Error loading state: state snapshot was created by Terraform v1.10.2, which
is newer than current v1.9.8; upgrade to Terraform v1.10.2 or greater to work with
this state

That is the classic "someone ran apply from a laptop with a newer Terraform and now CI is broken". The fix is to upgrade CI - there is no downgrade.

Upgrading Terraform: read the changelog (the list of what changed in each release). 1.x promises that configuration keeps working, but deprecations and new warnings appear. Bump required_version and the CI pin together, run init, validate and a plan in every environment - no changes expected - then merge.

Pinning providers

In the root configuration:

terraform {
  required_providers {
    azurerm = {
      source  = "hashicorp/azurerm"
      version = "~> 4.14"
    }
  }
}

plus the committed .terraform.lock.hcl, which records the exact version chosen (e.g. 4.14.0) and its checksums. ~> 4.14 allows 4.14 and later 4.x, but not 5.0.

Modules state a minimum only (>= 4.0): a configuration can load just one version of each provider, so if every module pinned tightly they would conflict. The root chooses.

Minor and patch upgrades

terraform init -upgrade          # newest version the constraints allow
git diff .terraform.lock.hcl     # what moved
terraform plan                   # every environment: expect no changes

Automate the boring part: Dependabot (GitHub's) and Renovate (a similar open-source bot) watch for new versions and open PRs that bump the constraint and the lock file. CI runs the plans; a human reads them.

Major upgrades (azurerm 3.x -> 4.0)

A major version may rename, remove or change defaults. The azurerm 4.0 upgrade guide (a page in the provider docs listing every breaking change) includes, among many others:

provider           subscription_id is required (or ARM_SUBSCRIPTION_ID)
                   skip_provider_registration -> resource_provider_registrations
storage account    enable_https_traffic_only  -> https_traffic_only_enabled
cluster (AKS)      automatic_channel_upgrade  -> automatic_upgrade_channel
                   node_os_channel_upgrade    -> node_os_upgrade_channel

(The last two are arguments of the cluster resource from 14.2: how its software and its machines' operating system get updates. Only the names changed.)

The procedure:

  1. Read the upgrade guide for the major version, top to bottom.
  2. Create a branch; widen the constraint (~> 4.0 or ~> 4.14); run terraform init -upgrade.
  3. Run terraform validate and plan. Renamed arguments fail loudly:
│ Error: Unsupported argument
│
│   on main.tf line 18, in resource "azurerm_storage_account" "logs":
│   18:   enable_https_traffic_only = true
│
│ An argument named "enable_https_traffic_only" is not expected here.
  1. Fix them. Plan every environment. The target is No changes: the same values under new names. Anything that wants to replace a resource is a migration you have not finished - read the guide's section for that resource.
  2. Merge constraint + lock file + fixes together. Upgrade dev first, prod last.

The lab enforces the renames above against the version recorded in the lock file, so 3.x names fail on 4.x and the other way round (simulator: only these renames are modelled).

Multi-platform lock files

The lock file records a checksum per platform - operating system plus CPU type, like darwin_arm64 (a Mac with Apple Silicon) or linux_amd64 (a normal Linux server). A lock file created on your Mac only contains the Mac checksums, so a Linux CI runner may fail to verify the provider it downloads. Record all the platforms you use, once:

terraform providers lock \
  -platform=linux_amd64 \
  -platform=darwin_arm64 \
  -platform=windows_amd64

terraform providers lock downloads the provider for each -platform listed and writes all their checksums into the lock file.

Modules

A checklist for any upgrade PR

What the errors look like

Raise the constraint and run plain init (without -upgrade): the lock file still says 3.117.1, and init tells you what to do.

│ Error: Failed to query available provider packages
│
│ Could not retrieve the list of available versions for provider
│ hashicorp/azurerm: locked provider registry.terraform.io/hashicorp/azurerm
│ 3.117.1 does not match configured version constraint ~> 4.14; must use
│ terraform init -upgrade to allow selection of new versions

After init -upgrade, the lock file diff is the record of what moved:

 provider "registry.terraform.io/hashicorp/azurerm" {
-  version     = "3.117.1"
-  constraints = "~> 3.117"
+  version     = "4.14.0"
+  constraints = "~> 4.14"
   hashes = [
-    "h1:...",
+    "h1:...",

Lines starting with - were removed, + added: the version went from 3.117.1 to 4.14.0, the constraint changed, and the checksums are new.

Then validate shows every renamed or removed argument at once - which is the work list for the upgrade PR.

Terraform CLI upgrades in practice

# .terraform-version (read by tfenv; mise and asdf have equivalents)
1.9.8

tfenv install          # installs the pinned version
terraform version      # Terraform v1.9.8 on darwin_arm64

.terraform-version is a one-line file with the version number. tfenv install (with no argument) installs the version that file names; terraform version prints which binary you are actually running and on which platform.

terraform {
  required_version = "~> 1.9.0"   # root: tight
}

To move to 1.12: bump .terraform-version, the CI pin and required_version together in one PR; run init, validate and plan in every environment (expect No changes); merge. The first apply with 1.12 upgrades that state's format for good - which is why every pipeline that touches a state must be on 1.12 before any of them applies with it.

Renovate and Dependabot for Terraform

Both bots understand required_providers, module version / ?ref= and the lock file. A good setup:

What you can now do:

Why it helps

Upgrades are routine platform work, and done carelessly they break everyone. Situations: a colleague applies from a laptop with Terraform 1.10, and now CI on 1.9 cannot read the state, with no downgrade possible. The azurerm 3-to-4 upgrade renames enable_https_traffic_only and automatic_channel_upgrade, and validate gives you the whole work list at once. Renovate opens a provider bump PR and the plan for prod shows a replacement, which you must treat as a migration bug. A Linux agent fails checksum verification because the lock file only has Mac hashes. Owning upgrades well is a sign of a senior platform engineer, and interviewers like the "upgrade across environments" question.

FAQ

Can I downgrade Terraform after a newer version wrote the state?

No. State is forward-only: a state written by 1.10 records that version, and a 1.9 binary refuses to load it ("state snapshot was created by Terraform v1.10.2, which is newer than current v1.9.8"). The fix is upgrading every pipeline and laptop that uses that state. That is why the Terraform version bump must reach every pipeline before any of them applies with the new version.

Why use ~> 1.9.0 for Terraform but ~> 4.14 for providers?

They are different bets. Pinning the Terraform binary to patch releases in the root (~> 1.9.0) keeps everyone on the same minor, since a minor bump writes state that older binaries cannot read. Providers use a minor-level constraint because the lock file already pins the exact version, and init -upgrade moves within the range deliberately. Reusable modules state only a minimum for both.

What does init -upgrade do exactly?

It re-resolves providers to the newest versions allowed by the constraints, ignoring the lock file's current selection, installs them and rewrites .terraform.lock.hcl; for registry modules it also moves to the newest version matching version. Plain init honours the lock and refuses a constraint the locked version no longer satisfies, telling you to use -upgrade.

Should Renovate or Dependabot auto-merge Terraform updates?

Not for anything that touches production. They are great at opening grouped PRs that bump provider constraints, module versions and the lock file, and the PR pipeline then plans every environment. A human reads those plans. Auto-merge is acceptable at most for patch updates of dev-only tooling.

How do I know what an azurerm major upgrade will break?

Read the provider's upgrade guide for that major version end to end: it lists renamed and removed arguments, changed defaults and behavioural changes per resource. Then let the tools find your instances: after init -upgrade, terraform validate reports every unsupported argument at once, and the plan per environment shows anything that would change or be replaced.

In an interview Junior

How do you pin versions in Terraform, and how do you upgrade them safely?

Three things get pinned, each in its own place:

Upgrading: one PR. terraform init -upgrade, read the lock-file diff, read the changelog or upgrade guide for a major version, fix what terraform validate reports, then plan every environment - the target is No changes. Roll out dev first, prod last. Renovate or Dependabot can open these PRs; a human reads the plans.

Also asked: What does ~> 4.14 allow, and what does ~> 4.14.0 allow? · Someone applied with a newer Terraform from a laptop and now CI cannot read state. What do you do? · Why do modules state a minimum provider version instead of a tight pin?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.