OnCallReady

Lesson 13.21 · Terraform: State & Modules · 19 min read

Drift, import, and the commands that replaced taint and refresh

In plain words

Imagine you arrange your room exactly as your drawing shows. While you are at school, your little brother moves the lamp and takes a chair. When you come home you compare the room with the drawing and put everything back. But sometimes your parents moved the lamp on purpose, because the socket moved, and then the right thing is to change the drawing, not the room.

Drift is the difference between what Terraform's state says and what really exists. Every plan starts by refreshing, and reports "Objects have changed outside of Terraform" before planning to undo it. terraform plan -refresh-only shows only the drift. For each change you decide: revert it, codify it in the code, or declare that something else owns it with lifecycle { ignore_changes = [...] }. -replace forces a rebuild on purpose.

The problem

At 2 a.m. during an incident, someone opened a firewall rule by hand in the portal (Azure's web console). Nobody wrote it down. The next morning's routine Terraform change quietly removes that rule again - and the incident comes back.

The real world and your code drifted apart, and Terraform "fixed" it in the wrong direction. This lesson is about noticing drift, deciding which side is right, and the commands for each answer.

What you need to know already:

Drift

Drift is any difference between what state says and what really exists. It happens because Terraform is rarely the only thing that can change infrastructure:

How Terraform sees it

Every plan starts by refreshing: reading each object in state back from Azure.

azurerm_resource_group.main: Refreshing state... [id=/subscriptions/.../rg-orders-prod]
azurerm_network_security_group.web: Refreshing state... [id=/subscriptions/.../nsg-orders-web]

(azurerm_network_security_group is an NSG: a set of firewall rules, 13.11.)

If something changed or disappeared, the plan opens with a note:

Note: Objects have changed outside of Terraform

Terraform detected the following changes made outside of Terraform since the
last "terraform apply" which may have affected this plan:

  # azurerm_network_security_group.web has been deleted
  - resource "azurerm_network_security_group" "web" {
      - id   = "/subscriptions/.../networkSecurityGroups/nsg-orders-web" -> null
        name = "nsg-orders-web"
    }

  # azurerm_resource_group.main has changed
  ~ resource "azurerm_resource_group" "main" {
        id   = "/subscriptions/.../resourceGroups/rg-orders-prod"
        name = "rg-orders-prod"
      ~ tags = {
          ~ "owner" = "team-orders" -> "team-payments"
        }
    }

Unless you have made equivalent changes to your configuration, or ignored the
relevant attributes using ignore_changes, the following plan may include
actions to undo or respond to these changes.

Read it like this:

After the note comes the plan itself, and it plans to undo the drift: recreate what was deleted, revert what was edited. Terraform's model is simple: the configuration is the truth. Your job is deciding whether it really is.

Since Terraform 1.2 the note lists only outside changes that may affect the plan. If your code already has the change, or ignores that attribute, it stays quiet.

Looking at drift on its own: -refresh-only

terraform plan -refresh-only

shows only what changed outside Terraform and proposes no changes:

This is a refresh-only plan, so Terraform will not take any actions to undo
these. If you were expecting these changes then you can apply this plan to
record the updated values in the Terraform state without changing any remote
objects.

terraform apply -refresh-only then writes what it found into state, without touching Azure.

That is only half a fix. The configuration still says the old value, so the next normal plan will want to revert it again. The complete sequence for an outside change that was right:

  1. look: terraform plan -refresh-only;
  2. change the code to match;
  3. terraform plan - now clean for that object.

The decision, per object

the change outside was WRONG            let the normal plan revert it, apply
   (someone "cleaned up" an NSG)

the change outside was RIGHT            put it in the configuration; plan is clean
   (payments took over the service)

something else legitimately OWNS it     lifecycle { ignore_changes = [...] }
   (Azure Policy tags, autoscaler counts)

the object was deleted and should       delete the block (or use a removed block);
   stay deleted                         the plan has nothing to do

Never let a pipeline -auto-approve a plan that opens with Objects have changed outside of Terraform without a human reading it. Half of all drift is a deliberate change the code has not caught up with yet.

Deleted objects: what "recreate" means

When the plan recreates a deleted object, it is a new object: a new ID and, for anything holding data (storage, databases, Key Vaults), empty. For an NSG that is harmless. For a storage account the data only comes back from backups or soft delete.

So prevent deletes in the first place:

ignore_changes

A lifecycle block inside a resource changes how Terraform treats it. ignore_changes is one of its settings:

resource "azurerm_resource_group" "main" {
  name     = "rg-orders-prod"
  location = "westeurope"
  tags     = local.tags

  lifecycle {
    ignore_changes = [tags]
  }
}

Replacing on purpose: -replace

Sometimes an object is broken in a way Terraform cannot see: a virtual machine (VM) whose disk is corrupt, a TLS certificate (9.15) that must be issued again. -replace=ADDR (12.26) makes the plan destroy and recreate that one object:

terraform plan -replace='azurerm_linux_virtual_machine.jump'
terraform apply -replace='azurerm_linux_virtual_machine.jump'
  # azurerm_linux_virtual_machine.jump will be replaced, as requested
-/+ resource "azurerm_linux_virtual_machine" "jump" {

as requested tells the reviewer this replacement was asked for, not caused by a code change. It is in the plan for review, and nothing happens until apply.

What happened to taint and refresh

Two older commands did the same jobs without showing you first:

Both still exist and are deprecated (kept for now, not to be used) for the same reason: they change state without showing you what they will do.

If you find an object marked tainted (after a failed create, or an old taint), the plan says so:

  # azurerm_linux_virtual_machine.jump is tainted, so must be replaced

and terraform untaint ADDR removes the mark if replacing is not wanted.

Detecting drift continuously

A plan only notices drift when someone runs it. So teams run one on a schedule:

terraform plan -detailed-exitcode -input=false -lock-timeout=5m
# 0 = in sync   2 = drift or pending changes   1 = error

-detailed-exitcode turns the result into an exit code (12.26); -input=false never waits for typing; -lock-timeout=5m waits for a busy lock (13.11).

A nightly job per state that opens a ticket on exit code 2 turns "we found out during the next deploy" into "we found out the next morning". HCP Terraform offers the same as a built-in feature.

The -refresh=false shortcut

terraform plan -refresh=false skips the refresh and plans against state as recorded. Faster on huge states, and useful when Azure is limiting how many requests you may make. But it plans against a possibly old picture: deleted objects are not recreated, edited ones are not reverted. Never make it the default.

What you can now do:

Why it helps

Incidents leave drift behind. During the night someone opens an NSG rule in the portal; the next morning's pipeline plan wants to remove it, and if the pipeline auto-approves, the fix is undone mid-incident. Knowing to read the "changed outside of Terraform" note, and to codify the change first, prevents a repeat outage. Azure Policy stamping tags makes every plan noisy until you use ignore_changes = [tags["CreatedOnDate"]] for just that key. An autoscaler's change to a machine counts node_count constantly. And a plan that recreates a deleted storage account creates an empty one, which is why prevention (resource locks, prevent_destroy) matters. The exam asks what replaced taint and refresh.

FAQ

What is the difference between plan -refresh-only and apply -refresh-only?

terraform plan -refresh-only shows what changed outside Terraform and proposes no actions. terraform apply -refresh-only writes those changes into state, without touching Azure. That is only half a fix: the configuration still describes the old value, so the next normal plan will want to revert it. For a correct outside change, update the code to match, then plan until it is clean.

Should I just add ignore_changes whenever a plan shows drift?

No. ignore_changes is for attributes another system legitimately owns: tags stamped by Azure Policy, a node count managed by an autoscaler. Using it to silence drift you should fix means the configuration stops describing reality. Target narrowly: ignore_changes = [tags["CreatedOnDate"]] ignores only that key, while ignore_changes = [tags] or all ignores far more than you mean.

Why were taint and refresh deprecated?

Both changed state immediately without a plan to review. taint marked an object for replacement; its successor -replace=ADDR (0.15.2) shows the replacement in a plan first. refresh rewrote state from reality; with wrong credentials or a misconfigured subscription it could record that everything was gone. apply -refresh-only shows the drift and asks before saving.

If someone deletes a storage account in the portal, will terraform apply bring it back?

It creates a new storage account with the same configuration: new object, and empty. The data is only recoverable from backups or Azure soft delete, not from Terraform. Prevention is the real answer: CanNotDelete resource locks on production, portal write access only through just-in-time elevation, and prevent_destroy for destroys that Terraform itself would plan.

Is -refresh=false a good way to speed up plans?

Only occasionally, for example when an API is rate-limiting you on a huge state. It plans against the state as recorded, so deleted objects will not be recreated and edited ones will not be reverted; you are planning against a possibly stale picture. Never make it the default. The real fix for slow plans is usually splitting the state.

In an interview Junior

What is configuration drift in Terraform, and how do you detect and handle it?

Drift is any difference between state and reality: someone changed a firewall rule by hand during an incident, a policy added tags, an autoscaler changed a count.

Detect: every plan refreshes first, and drift appears as Objects have changed outside of Terraform above the plan. terraform plan -refresh-only shows only the drift. A nightly plan -detailed-exitcode (exit 2 = changes) finds it without waiting for the next deploy.

Handle, per object:

Never let a pipeline auto-approve a plan that opens with that drift note: half of drift is a deliberate fix the code has not caught up with.

Also asked: What is the difference between terraform plan -refresh-only and a normal plan? · What replaced terraform taint, and why? · What does ignore_changes do, and when is it the wrong answer?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.