The problem
At 2 a.m. during an incident, someone opened a firewall rule by hand in the portal (Azure's web console). Nobody wrote it down. The next morning's routine Terraform change quietly removes that rule again - and the incident comes back.
The real world and your code drifted apart, and Terraform "fixed" it in the wrong direction. This lesson is about noticing drift, deciding which side is right, and the commands for each answer.
What you need to know already:
- state and refresh: every plan re-reads objects from Azure (13.1, 12.24)
- plan modes
-refresh-only,-replace,-detailed-exitcode(12.26) - this lesson uses them for real - reading a plan:
~change,-/+replace (12.24)
Drift
Drift is any difference between what state says and what really exists. It happens because Terraform is rarely the only thing that can change infrastructure:
- someone clicks in the portal during an incident;
- Azure Policy (rules Azure applies to resources automatically, such as adding required tags) adds tags or settings on a schedule;
- an autoscaler (a service that adds or removes machines as load changes) changes a machine count;
- another tool (a script, a second Terraform state) manages the same object;
- Azure itself changes a default or retires a SKU (a product size or tier, like a VM size).
How Terraform sees it
Every plan starts by refreshing: reading each object in state back from Azure.
azurerm_resource_group.main: Refreshing state... [id=/subscriptions/.../rg-orders-prod]
azurerm_network_security_group.web: Refreshing state... [id=/subscriptions/.../nsg-orders-web]
(azurerm_network_security_group is an NSG: a set of firewall rules, 13.11.)
If something changed or disappeared, the plan opens with a note:
Note: Objects have changed outside of Terraform
Terraform detected the following changes made outside of Terraform since the
last "terraform apply" which may have affected this plan:
# azurerm_network_security_group.web has been deleted
- resource "azurerm_network_security_group" "web" {
- id = "/subscriptions/.../networkSecurityGroups/nsg-orders-web" -> null
name = "nsg-orders-web"
}
# azurerm_resource_group.main has changed
~ resource "azurerm_resource_group" "main" {
id = "/subscriptions/.../resourceGroups/rg-orders-prod"
name = "rg-orders-prod"
~ tags = {
~ "owner" = "team-orders" -> "team-payments"
}
}
Unless you have made equivalent changes to your configuration, or ignored the
relevant attributes using ignore_changes, the following plan may include
actions to undo or respond to these changes.
Read it like this:
has been deletedwithid = ... -> null: the NSG is gone from Azure.has changedwith~ "owner" = "team-orders" -> "team-payments": someone edited the tag; the arrow goes from the value in state to the value found in Azure.
After the note comes the plan itself, and it plans to undo the drift: recreate what was deleted, revert what was edited. Terraform's model is simple: the configuration is the truth. Your job is deciding whether it really is.
Since Terraform 1.2 the note lists only outside changes that may affect the plan. If your code already has the change, or ignores that attribute, it stays quiet.
Looking at drift on its own: -refresh-only
terraform plan -refresh-only
shows only what changed outside Terraform and proposes no changes:
This is a refresh-only plan, so Terraform will not take any actions to undo
these. If you were expecting these changes then you can apply this plan to
record the updated values in the Terraform state without changing any remote
objects.
terraform apply -refresh-only then writes what it found into state, without touching Azure.
That is only half a fix. The configuration still says the old value, so the next normal plan will want to revert it again. The complete sequence for an outside change that was right:
- look:
terraform plan -refresh-only; - change the code to match;
terraform plan- now clean for that object.
The decision, per object
the change outside was WRONG let the normal plan revert it, apply
(someone "cleaned up" an NSG)
the change outside was RIGHT put it in the configuration; plan is clean
(payments took over the service)
something else legitimately OWNS it lifecycle { ignore_changes = [...] }
(Azure Policy tags, autoscaler counts)
the object was deleted and should delete the block (or use a removed block);
stay deleted the plan has nothing to do
Never let a pipeline -auto-approve a plan that opens with Objects have changed outside of Terraform without a human reading it. Half of all drift is a deliberate change the code has not caught up with yet.
Deleted objects: what "recreate" means
When the plan recreates a deleted object, it is a new object: a new ID and, for anything holding data (storage, databases, Key Vaults), empty. For an NSG that is harmless. For a storage account the data only comes back from backups or soft delete.
So prevent deletes in the first place:
- Azure resource locks (
CanNotDelete, 13.4) on production resource groups; - write access in the portal only as an exception, granted for an hour when needed ("just-in-time" access);
lifecycle { prevent_destroy = true }on the Terraform side, so a Terraform plan that would destroy it fails (13.46).
ignore_changes
A lifecycle block inside a resource changes how Terraform treats it. ignore_changes is one of its settings:
resource "azurerm_resource_group" "main" {
name = "rg-orders-prod"
location = "westeurope"
tags = local.tags
lifecycle {
ignore_changes = [tags]
}
}
- The listed attributes are ignored when comparing - Terraform will not revert changes to them. It still sets them when it creates the object.
ignore_changes = allignores every attribute: "create it and never touch it again". Rarely what you want.- You can target one map key:
ignore_changes = [tags["CreatedOnDate"]]. Only that key is ignored, so Azure Policy can stamp its date and your own tags are still enforced. Almost always better than ignoring alltags. - It is for attributes another system owns. Using it to silence drift you should fix is how a configuration stops describing reality.
Replacing on purpose: -replace
Sometimes an object is broken in a way Terraform cannot see: a virtual machine (VM) whose disk is corrupt, a TLS certificate (9.15) that must be issued again. -replace=ADDR (12.26) makes the plan destroy and recreate that one object:
terraform plan -replace='azurerm_linux_virtual_machine.jump'
terraform apply -replace='azurerm_linux_virtual_machine.jump'
# azurerm_linux_virtual_machine.jump will be replaced, as requested
-/+ resource "azurerm_linux_virtual_machine" "jump" {
as requested tells the reviewer this replacement was asked for, not caused by a code change. It is in the plan for review, and nothing happens until apply.
What happened to taint and refresh
Two older commands did the same jobs without showing you first:
terraform taint ADDRmarked an object in state for replacement - writing to state immediately, with no plan to review. Replaced by-replace=ADDR(Terraform 0.15.2).terraform refreshwrote refreshed values into state with no review step. With the wrong credentials it could "discover" that every resource was gone. Replaced byapply -refresh-only, which shows the drift and asks.
Both still exist and are deprecated (kept for now, not to be used) for the same reason: they change state without showing you what they will do.
If you find an object marked tainted (after a failed create, or an old taint), the plan says so:
# azurerm_linux_virtual_machine.jump is tainted, so must be replaced
and terraform untaint ADDR removes the mark if replacing is not wanted.
Detecting drift continuously
A plan only notices drift when someone runs it. So teams run one on a schedule:
terraform plan -detailed-exitcode -input=false -lock-timeout=5m
# 0 = in sync 2 = drift or pending changes 1 = error
-detailed-exitcode turns the result into an exit code (12.26); -input=false never waits for typing; -lock-timeout=5m waits for a busy lock (13.11).
A nightly job per state that opens a ticket on exit code 2 turns "we found out during the next deploy" into "we found out the next morning". HCP Terraform offers the same as a built-in feature.
The -refresh=false shortcut
terraform plan -refresh=false skips the refresh and plans against state as recorded. Faster on huge states, and useful when Azure is limiting how many requests you may make. But it plans against a possibly old picture: deleted objects are not recreated, edited ones are not reverted. Never make it the default.
What you can now do:
- read the Objects have changed outside of Terraform note
- decide per object: revert, codify, ignore, or let it go
- use
-refresh-only,ignore_changesand-replacefor what each is for