The infrastructure pipeline has failed four times since 06:15. Nobody can deploy. Every run, and every plan you try yourself, ends the same way:
$ terraform plan
Acquiring state lock. This may take a few moments...
╷
│ Error: Error acquiring the state lock
│
│ Error message: state blob is already locked
│ Lock Info:
│ ID: 6f0c1a2b-9d3e-4c5f-8a7b-e1d2c3b4a596
│ Path: tfstate/orders/dev.tfstate
│ Operation: OperationTypeApply
│ Who: AzDevOps@fv-az412-118
│ Version: 1.9.8
│ Created: 2026-09-23 06:12:44.518417 +0000 UTC
│ Info:
│
│
│ Terraform acquires a state lock to protect the state from being written
│ by multiple users at the same time. Please resolve the issue above and try
│ again. For most commands, you can disable locking with the "-lock=false"
│ flag, but this is not recommended.
╵The tempting fix is in the error message itself. Do not reach for -lock=false, and do not force-unlock yet.
What the lock is for
Terraform's state is the record of what it manages. Every command that might write it (plan, apply, destroy, import, state mv/rm/push) takes a lock first and releases it at the end. Without it, two runs can read the same state, each make changes, and the last one to write silently erases the other's record. The resources still exist in the cloud, but state forgot them.
Where the lock lives depends on the backend: a lease on the state blob for azurerm, a .tflock object next to the state for S3 (Terraform 1.10+, or a DynamoDB table before that), built in for GCS and HCP Terraform.
So the error is the system working. The only question is whether the run that holds the lock is still alive.
The diagnosis path
1. Read the Lock Info
- Who:
AzDevOps@fv-az412-118- a CI agent, not a laptop. - Created: 06:12, before the first failed run at 06:15.
- Operation:
OperationTypeApply- an apply, so it may have changed real infrastructure. - Path: which state. Check it is the one you meant.
2. Find the holder
$ ls -l /srv/ci/runs/
total 24
-rw-r--r-- 1 root root 355 Sep 22 20:00 4126.log
-rw-r--r-- 1 root root 564 Sep 22 20:00 4127.log
-rw-r--r-- 1 root root 481 Sep 22 20:00 4128.log
...
$ cat /srv/ci/runs/4127.log
run: 4127
pipeline: orders-infra
branch: main (PR #212 merged)
agent: fv-az412-118
started: 2026-09-23T06:12:31Z
finished: 2026-09-23T06:14:02Z
result: canceled
terraform apply tfplan
Acquiring state lock. This may take a few moments...
azurerm_storage_account.exports: Creating...
azurerm_storage_account.exports: Still creating... [10s elapsed]
azurerm_storage_account.exports: Still creating... [20s elapsed]
##[error]The operation was canceled.
##[section]Finishing: terraform applyRun 4127, on that agent, started 13 seconds before the lock's timestamp, and was cancelled in the middle of an apply. The agent went away; the lease on the blob did not.
3. Decide: wait or unlock
If the holder were still running - a long apply, a colleague in the middle of a change - the answer is to wait. Let Terraform do the waiting for you:
terraform plan -lock-timeout=10mUnlocking a live run causes exactly the overwrite that locking exists to prevent. Here the run is provably dead: cancelled, with a finish time.
4. Force-unlock with that exact ID
$ terraform force-unlock 6f0c1a2b-9d3e-4c5f-8a7b-e1d2c3b4a596
Do you really want to force-unlock?
Terraform will remove the lock on the remote state.
This will allow local Terraform commands to modify this state, even though it
may still be in use. Only 'yes' will be accepted to confirm.
Enter a value: yes
Terraform state has been successfully unlocked!
The state has been unlocked, and Terraform commands should now be able to
obtain a new lock on the remote state.The ID protects you: if someone took a new lock after you read the old one, the IDs no longer match and Terraform refuses.
The part that bites later: what the dead run left behind
A cancelled apply does not roll back. Run 4127 was creating a storage account when it died. Azure finished creating it; Terraform never wrote it to state. The next apply tries to create it again:
$ terraform apply -auto-approve
...
azurerm_storage_account.exports: Creating...
╷
│ Error: A resource with the ID "/subscriptions/00000000-1111-2222-3333-444444444444/resourceGroups/rg-orders-dev/providers/Microsoft.Storage/storageAccounts/stordersexportsdev" already exists - to be managed via Terraform this resource needs to be imported into the State. Please see the resource documentation for "azurerm_storage_account" for more information.
╵That object is an orphan: real, but unknown to state. Adopt it, do not delete it - it may already hold data:
$ terraform import azurerm_storage_account.exports /subscriptions/00000000-1111-2222-3333-444444444444/resourceGroups/rg-orders-dev/providers/Microsoft.Storage/storageAccounts/stordersexportsdev
...
Import successful!
$ terraform plan
...
No changes. Your infrastructure matches the configuration.The address comes from main.tf, the Azure ID from the error message. On Terraform 1.5+ you can do the same reviewably with an import block in the code. A clean plan means state matches reality again, and the pipeline can run.
Keeping it from coming back
- One run per state at a time. Use the CI system's own queueing: Azure DevOps exclusive locks on the environment, or a GitHub Actions
concurrencygroup per environment withcancel-in-progress: false. - Approve before the apply stage starts, not while it runs. Cancelling an apply half-way is how stale locks and orphans are made.
-lock-timeout=5mor so in pipelines, so two runs a few seconds apart queue instead of failing.- A written unlock procedure: Lock Info, find the holder, prove it is gone, unlock with that ID, then plan and look for orphans.
-lock=falseis not on the list.
Locks only protect one state from concurrent writers. They do nothing when a directory reads the wrong state, or when a list change re-addresses resources. And the state you just unlocked holds every secret your resources have, in plain text - a reason to keep its storage and the pipeline's logs locked down: a password in a CI log.
Practise it
Incident: the state is locked and the pipeline is red (13.12) has the stale lock, the run logs and the orphan. Drill: the state is locked - wait or unlock? (13.14) varies the holder, so you practise the decision, not just the command.