The problem locking solves
Your pipeline and your colleague both run terraform apply against the same state at the same moment. Neither sees an error. Three days later a plan wants to create a storage account that has existed all along. What happened?
Here is the timeline:
10:00:01 run A reads state, serial 41
10:00:03 run B reads state, serial 41
10:00:40 run A writes serial 42 (adds the storage account)
10:00:55 run B writes serial 42 (adds the NSG) - A's storage account is gone from state
(An NSG, network security group, is a set of Azure firewall rules for a subnet.)
Run B started from the old state and wrote its result over A's. The storage account still exists in Azure, but state no longer mentions it. The next plan wants to create it, and the create fails with already exists.
State locking stops this: the second run waits or fails instead of writing.
What you need to know already:
- serial, and why two runs writing at once lose data (13.1)
- backends and the azurerm blob lease (13.4)
- saved plans with
plan -out(12.24)
How it looks
Every command that could write state takes the lock first and releases it at the end:
$ terraform plan
Acquiring state lock. This may take a few moments...
azurerm_resource_group.orders: Refreshing state... [id=/subscriptions/.../rg-orders-dev]
...
Releasing state lock. This may take a few moments...
The middle lines are the refresh: Terraform re-reading each object from Azure (12.24). The lock is held for the whole run.
Commands that lock: plan, apply, destroy, import, refresh, state mv, state rm, state push, taint/untaint (old commands, 13.21), and init when it moves state. Commands that only read and do not lock: state list, state show, state pull, output, show.
Where the lock lives depends on the backend:
- local backend: an operating-system file lock. It only stops other programs on the same machine, so it cannot protect a team.
- azurerm: a lease on the state blob (13.4).
- s3: a
.tflockfile next to the state (1.10+), or an older database table (DynamoDB). - gcs and HCP Terraform: built in.
When the lock is taken
This is what you see when someone else holds the lock:
# while the pipeline's run holds the lock (the lock incident)
terraform apply
Acquiring state lock. This may take a few moments...
╷
│ Error: Error acquiring the state lock
│
│ Error message: state blob is already locked
│ Lock Info:
│ ID: 6f0c1a2b-9d3e-4c5f-8a7b-e1d2c3b4a596
│ Path: tfstate/orders/dev.tfstate
│ Operation: OperationTypeApply
│ Who: AzDevOps@fv-az412-118
│ Version: 1.9.8
│ Created: 2026-09-23 06:12:44.518417 +0000 UTC
│ Info:
│
│
│ Terraform acquires a state lock to protect the state from being written
│ by multiple users at the same time. Please resolve the issue above and try
│ again. For most commands, you can disable locking with the "-lock=false"
│ flag, but this is not recommended.
╵
The Lock Info block tells you everything you need to decide what to do:
ID the lock's identity - force-unlock needs exactly this value
Path which state (container/key) - is it even the state you meant?
Operation OperationTypeApply / OperationTypePlan / OperationTypeInvalid (state commands)
Who user@host that took it - a pipeline agent's name, or a laptop
Version the Terraform version of the holder
Created when - a lock from two minutes ago and one from last night are different stories
AzDevOps@fv-az412-118 is a pipeline agent: the machine a CI system (here Azure DevOps, Microsoft's CI service) runs a job on. A laptop would show something like alice@alice-mbp.
That error is the system working. Another run is (or was) in the middle of something.
Waiting instead of failing: -lock-timeout
In CI, two pipeline runs a few seconds apart are normal. Rather than fail the second one at once, let it wait for the lock:
terraform plan -lock-timeout=5m
terraform apply -lock-timeout=5m tfplan
-lock-timeout=5m means: keep retrying for up to five minutes, then give up. Five to ten minutes is typical.
If runs queue longer than that, the pipeline itself should allow only one run per state at a time. CI systems have a setting for it (Azure DevOps "exclusive locks", GitHub Actions concurrency: groups, shown at the end).
When the holder is dead: force-unlock
A pipeline that is cancelled, or an agent that crashes in the middle of an apply, can leave the lease behind. Every later run fails on a lock that nobody is using any more.
terraform force-unlock <ID> removes it. The procedure, in this order, every time:
- Read the Lock Info. Who, since when, which operation.
- Find the holder. The pipeline run on that agent, the colleague on that laptop.
- Confirm it is really gone. The run shows cancelled or failed, the agent is gone, the colleague has closed their terminal. If it is still running - wait. Unlocking a live run causes exactly the damage locking exists to prevent.
- Unlock with that exact ID:
# only with the ID from that error, once the holder is confirmed dead
terraform force-unlock 6f0c1a2b-9d3e-4c5f-8a7b-e1d2c3b4a596
Do you really want to force-unlock?
Terraform will remove the lock on the remote state.
This will allow local Terraform commands to modify this state, even though it
may still be in use. Only 'yes' will be accepted to confirm.
Enter a value: yes
Terraform state has been successfully unlocked!
A wrong ID is refused. That protects you from removing a new lock that someone took after you read the old one:
│ Error: Failed to unlock state: lock ID "0000..." does not match existing lock ID "6f0c..."
- Assume the dead run did some of its work. An apply that died after creating something but before writing state leaves an orphan: it exists in Azure but not in state. The next apply fails with A resource with the ID ... already exists - to be managed via Terraform this resource needs to be imported into the State. Import it (13.25); do not delete it (it may already hold data).
terraform force-unlock -force <ID> skips the yes/no prompt. That is for a written procedure after a human has done steps 1-3, not for pipelines.
-lock=false
Every command that locks accepts -lock=false: run without taking the lock. On shared state, that is choosing the corruption above on purpose.
Legitimate uses are rare: a local backend on a machine where you know nothing else runs, or a backend that cannot lock at all.
What a lock does NOT protect against
- Two different states managing the same object. Each state has its own lock, so both will happily fight over one resource group. That is a design error, not a locking problem.
- Changes made in the portal (Azure's web console). Locks belong to Terraform; Azure knows nothing about them. Those changes are drift (13.21); Azure resource locks and roles are the defence.
- A stale plan. A lock is held only while a command runs. A plan saved an hour ago is protected by the serial check (Saved plan is stale, 13.1), not by a lock.
- Hand edits to the blob. Anyone with Blob Data Contributor can overwrite the state blob directly. Give that role only to pipeline identities.
In a pipeline
A pipeline is defined in a YAML file in the repo. In GitHub Actions (GitHub's CI service), a concurrency group allows one run at a time per group name:
# GitHub Actions: one run at a time per environment's state
concurrency:
group: terraform-${{ inputs.env }}
cancel-in-progress: false # never cancel a running apply
group names the queue (one per environment here). cancel-in-progress: false matters: cancelling an apply half-way is how stale locks and orphans are made. Approve plans before the apply stage starts, not while it is running.
What you can now do:
- read Lock Info and say who holds the lock, since when, doing what
- choose between waiting (
-lock-timeout) andforce-unlock, and follow the unlock procedure - explain what a lock does not protect against