OnCallReady

Lesson 13.11 · Terraform: State & Modules · 18 min read

Locking: what it prevents, and how to handle a held lock

In plain words

Imagine two kids editing the same drawing on paper at the same time. Each takes a copy home, adds something, and brings it back. Whoever comes back second replaces the drawing with theirs, and the first kid's work silently disappears. The fix is a talking stick: only the kid holding it may change the drawing, everyone else waits.

A state lock is that talking stick. On the azurerm backend it is a lease on the state blob, taken by plan, apply and the state commands and released afterwards. When someone else has it, you get "Error acquiring the state lock" with the holder's name and a lock ID. If the holder crashed and the stick is left lying around, terraform force-unlock <ID> puts it back, but only after you are sure the holder is really gone.

The problem locking solves

Your pipeline and your colleague both run terraform apply against the same state at the same moment. Neither sees an error. Three days later a plan wants to create a storage account that has existed all along. What happened?

Here is the timeline:

10:00:01  run A reads state, serial 41
10:00:03  run B reads state, serial 41
10:00:40  run A writes serial 42  (adds the storage account)
10:00:55  run B writes serial 42  (adds the NSG) - A's storage account is gone from state

(An NSG, network security group, is a set of Azure firewall rules for a subnet.)

Run B started from the old state and wrote its result over A's. The storage account still exists in Azure, but state no longer mentions it. The next plan wants to create it, and the create fails with already exists.

State locking stops this: the second run waits or fails instead of writing.

What you need to know already:

How it looks

Every command that could write state takes the lock first and releases it at the end:

$ terraform plan
Acquiring state lock. This may take a few moments...
azurerm_resource_group.orders: Refreshing state... [id=/subscriptions/.../rg-orders-dev]
...
Releasing state lock. This may take a few moments...

The middle lines are the refresh: Terraform re-reading each object from Azure (12.24). The lock is held for the whole run.

Commands that lock: plan, apply, destroy, import, refresh, state mv, state rm, state push, taint/untaint (old commands, 13.21), and init when it moves state. Commands that only read and do not lock: state list, state show, state pull, output, show.

Where the lock lives depends on the backend:

When the lock is taken

This is what you see when someone else holds the lock:

# while the pipeline's run holds the lock (the lock incident)
terraform apply
Acquiring state lock. This may take a few moments...
╷
│ Error: Error acquiring the state lock
│
│ Error message: state blob is already locked
│ Lock Info:
│   ID:        6f0c1a2b-9d3e-4c5f-8a7b-e1d2c3b4a596
│   Path:      tfstate/orders/dev.tfstate
│   Operation: OperationTypeApply
│   Who:       AzDevOps@fv-az412-118
│   Version:   1.9.8
│   Created:   2026-09-23 06:12:44.518417 +0000 UTC
│   Info:
│
│
│ Terraform acquires a state lock to protect the state from being written
│ by multiple users at the same time. Please resolve the issue above and try
│ again. For most commands, you can disable locking with the "-lock=false"
│ flag, but this is not recommended.
╵

The Lock Info block tells you everything you need to decide what to do:

ID         the lock's identity - force-unlock needs exactly this value
Path       which state (container/key) - is it even the state you meant?
Operation  OperationTypeApply / OperationTypePlan / OperationTypeInvalid (state commands)
Who        user@host that took it - a pipeline agent's name, or a laptop
Version    the Terraform version of the holder
Created    when - a lock from two minutes ago and one from last night are different stories

AzDevOps@fv-az412-118 is a pipeline agent: the machine a CI system (here Azure DevOps, Microsoft's CI service) runs a job on. A laptop would show something like alice@alice-mbp.

That error is the system working. Another run is (or was) in the middle of something.

Waiting instead of failing: -lock-timeout

In CI, two pipeline runs a few seconds apart are normal. Rather than fail the second one at once, let it wait for the lock:

terraform plan -lock-timeout=5m
terraform apply -lock-timeout=5m tfplan

-lock-timeout=5m means: keep retrying for up to five minutes, then give up. Five to ten minutes is typical.

If runs queue longer than that, the pipeline itself should allow only one run per state at a time. CI systems have a setting for it (Azure DevOps "exclusive locks", GitHub Actions concurrency: groups, shown at the end).

When the holder is dead: force-unlock

A pipeline that is cancelled, or an agent that crashes in the middle of an apply, can leave the lease behind. Every later run fails on a lock that nobody is using any more.

terraform force-unlock <ID> removes it. The procedure, in this order, every time:

  1. Read the Lock Info. Who, since when, which operation.
  2. Find the holder. The pipeline run on that agent, the colleague on that laptop.
  3. Confirm it is really gone. The run shows cancelled or failed, the agent is gone, the colleague has closed their terminal. If it is still running - wait. Unlocking a live run causes exactly the damage locking exists to prevent.
  4. Unlock with that exact ID:
# only with the ID from that error, once the holder is confirmed dead
terraform force-unlock 6f0c1a2b-9d3e-4c5f-8a7b-e1d2c3b4a596
Do you really want to force-unlock?
  Terraform will remove the lock on the remote state.
  This will allow local Terraform commands to modify this state, even though it
  may still be in use. Only 'yes' will be accepted to confirm.

  Enter a value: yes

Terraform state has been successfully unlocked!

A wrong ID is refused. That protects you from removing a new lock that someone took after you read the old one:

│ Error: Failed to unlock state: lock ID "0000..." does not match existing lock ID "6f0c..."
  1. Assume the dead run did some of its work. An apply that died after creating something but before writing state leaves an orphan: it exists in Azure but not in state. The next apply fails with A resource with the ID ... already exists - to be managed via Terraform this resource needs to be imported into the State. Import it (13.25); do not delete it (it may already hold data).

terraform force-unlock -force <ID> skips the yes/no prompt. That is for a written procedure after a human has done steps 1-3, not for pipelines.

-lock=false

Every command that locks accepts -lock=false: run without taking the lock. On shared state, that is choosing the corruption above on purpose.

Legitimate uses are rare: a local backend on a machine where you know nothing else runs, or a backend that cannot lock at all.

What a lock does NOT protect against

In a pipeline

A pipeline is defined in a YAML file in the repo. In GitHub Actions (GitHub's CI service), a concurrency group allows one run at a time per group name:

# GitHub Actions: one run at a time per environment's state
concurrency:
  group: terraform-${{ inputs.env }}
  cancel-in-progress: false      # never cancel a running apply

group names the queue (one per environment here). cancel-in-progress: false matters: cancelling an apply half-way is how stale locks and orphans are made. Approve plans before the apply stage starts, not while it is running.

What you can now do:

Why it helps

The 9am scenario: every Terraform pipeline for the orders environment fails with "state blob is already locked", Who is an Azure DevOps agent, Created is last night. Knowing the procedure (read Lock Info, find the run, confirm it is dead, force-unlock with the exact ID, then check for orphans the dead apply left behind) turns panic into a ten-minute fix. The opposite mistake is worse: force-unlocking a run that is still applying, which produces exactly the lost-update corruption locks exist to prevent. You will also add -lock-timeout=5m and pipeline concurrency groups so overlapping runs wait instead of failing. The exam asks which commands lock and what force-unlock needs.

Commands in this lesson

terraform

FAQ

Which commands take the state lock?

Anything that could write state: plan, apply, destroy, import, refresh, state mv, state rm, state push, taint and untaint, and init when it migrates state. Read-only commands such as state list, state show, state pull, output and show do not lock. Plan locks because its refresh must read a consistent state.

Does the local backend lock?

Only against other processes on the same machine, using an OS file lock. It cannot protect a team, because a colleague's laptop or a pipeline agent has its own copy of the file. Remote backends lock for everyone: azurerm with a blob lease, s3 with a lock object (1.10+) or the older DynamoDB table, gcs natively, HCP Terraform natively.

Why does force-unlock need the lock ID?

So you unlock exactly the lock you investigated. If the stale lock was cleared and a new run took a fresh lock in the meantime, the IDs differ and Terraform refuses: "lock ID does not match existing lock ID". Without that check you could unlock a live run you never looked at. -force skips the confirmation prompt, not the ID check.

A pipeline crashed mid-apply. What else should I worry about after unlocking?

Orphans. An apply that died after creating an object but before writing state leaves something that exists in Azure but not in state. The next apply fails with "A resource with the ID ... already exists - to be managed via Terraform this resource needs to be imported". Import it rather than deleting it, since it may already hold data or be referenced.

Does locking protect me from changes made in the portal?

No. Locks are Terraform's own mechanism and Azure knows nothing about them. Portal edits are drift, detected by the next plan's refresh; Azure RBAC and resource locks are the defence. Locks also do not protect against two different states managing the same object, or against hand edits of the blob by someone with write access.

In an interview Junior

A Terraform pipeline fails with "Error acquiring the state lock". What do you do?

That error is locking working: another run holds the lock on this state. Never just use -lock=false.

  1. Read the Lock Info: ID, Path (which state), Who (a pipeline agent or a laptop), Operation, Created.
  2. Find the holder - that pipeline run or that colleague.
  3. If it is still running, wait. In pipelines, -lock-timeout=5m makes the second run wait instead of failing, and the CI setting for one run per state at a time avoids the queue.
  4. If the holder is confirmed dead (cancelled run, crashed agent): terraform force-unlock <ID> with exactly that ID.
  5. Assume the dead run did part of its work: an object it created before dying may be an orphan (it exists but is not in state). Import it; do not delete it.

A lock does not protect against two states managing the same object, changes made by hand, or a stale saved plan.

Also asked: What is state locking and why does it matter? · Which Terraform commands take the state lock? · What can go wrong if you force-unlock a lock whose run is still alive?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.