OnCallReady

Lesson 12.20 · Terraform: Language & Workflow · 27 min read

count vs for_each - the one that causes outages

In plain words

Imagine a classroom where the teacher knows children only by their seat number. If the child in seat 2 moves away, everyone behind shuffles forward one seat, and the teacher now thinks every one of them is a different child. Chaos. In a classroom where the teacher knows children by name, one child leaving changes nothing for anyone else.

count is the seat-number classroom: instances are azurerm_subnet.s[0], [1], [2]. Remove one from the middle of the list and every instance after it gets a new identity, which for subnets and VMs means destroy and recreate. for_each is the name classroom: azurerm_subnet.s["web"], ["app"]. Remove "app" and only "app" is touched.

Why this matters

What you need to know already: addresses (12.18), maps, sets and lists (12.8), for expressions (12.11), flatten and toset (12.14), subnets inside a VNet (12.8).

Most real configurations create several of the same thing: five subnets, three machines. Terraform has two ways to say "make N of these": count and for_each. They look interchangeable. They are not - picking the wrong one is how "remove one subnet" turns into "destroy four".

Each copy is called an instance of the resource.

The difference is how instances are ADDRESSED

# count -> addressed by INDEX
resource "azurerm_subnet" "s" {
  count = 3
  name  = var.subnets[count.index]
  # ...
}
#  azurerm_subnet.s[0]  azurerm_subnet.s[1]  azurerm_subnet.s[2]
# for_each -> addressed by KEY
resource "azurerm_subnet" "s" {
  for_each = { web = "10.0.1.0/24", app = "10.0.2.0/24", db = "10.0.3.0/24" }
  name     = each.key
  # ...
}
#  azurerm_subnet.s["web"]  azurerm_subnet.s["app"]  azurerm_subnet.s["db"]

That is the whole difference, and it is enormous: the address is the identity Terraform uses to match configuration to state. Change an instance's address and Terraform sees one object removed and a different one added.

The production failure

Five subnets under count, in a list. Someone removes the second one. Terraform's view:

before                      after
s[0] = web                  s[0] = web       unchanged
s[1] = app       <-removed  s[1] = db        CHANGED: app -> db
s[2] = db                   s[2] = cache     CHANGED: db  -> cache
s[3] = cache                s[3] = mgmt      CHANGED: cache -> mgmt
s[4] = mgmt                 (gone)           DESTROYED

Everything after the removed element shifts down - exactly like removing item 1 from a JS array. A subnet's name cannot be changed in place, so each of those "changes" is a destroy and recreate:

  # azurerm_subnet.s[1] must be replaced
-/+ resource "azurerm_subnet" "s" {
      ~ id   = "/subscriptions/.../subnets/app" -> (known after apply)
      ~ name = "app" -> "db" # forces replacement
    }
...
Plan: 3 to add, 0 to change, 4 to destroy.

How to read it: -/+ means replace (destroy, then create); ~ marks an attribute that changes, old -> new; # forces replacement names the attribute that cannot change in place. The last line counts every action.

You removed one subnet and Terraform destroys four. Now imagine those are virtual machines, or databases. With for_each, removing "app" produces exactly one action - destroy s["app"] - because no other key's address changed.

The rule

Use for_each for anything with an identity. Use count for N identical anonymous copies, and for the on/off switch.

count is right for:

# a feature flag: zero or one (a bastion is a jump host you SSH through)
resource "azurerm_public_ip" "bastion" {
  count = var.enable_bastion ? 1 : 0
  # ...
}

# genuinely interchangeable copies
resource "azurerm_linux_virtual_machine" "worker" {
  count = var.worker_count
  name  = "vm-worker-${count.index}"
  # ...
}

Even the "identical workers" case is a trap if you ever remove one from the middle, so many teams use for_each there too, keyed by a name.

Referencing a conditional resource elsewhere needs care, because it is a list that may be empty:

output "bastion_ip" {
  value = one(azurerm_public_ip.bastion[*].ip_address)   # null when disabled
}

azurerm_public_ip.bastion[0].ip_address would fail when the count is 0 (there is no element 0 in an empty list).

What for_each accepts

A map, or a set of strings. Not a list:

╷
│ Error: Invalid for_each argument
│
│   on main.tf line 12, in resource "azurerm_subnet" "s":
│   12:   for_each = var.subnet_names
│
│ The given "for_each" argument value is unsuitable: the "for_each" argument
│ must be a map, or set of strings, and you have provided a value of type list
│ of string.
╵

Convert: for_each = toset(var.subnet_names). For a set, each.key and each.value are the same string.

A list of objects cannot become a set of strings. Turn it into a map with a for expression, choosing a unique, stable attribute as the key:

variable "subnets" {
  type = list(object({ name = string, cidr = string }))
}

resource "azurerm_subnet" "s" {
  for_each             = { for s in var.subnets : s.name => s }
  name                 = each.key
  address_prefixes     = [each.value.cidr]
  resource_group_name  = azurerm_resource_group.main.name
  virtual_network_name = azurerm_virtual_network.main.name
}

Choosing the key is the design decision. It must be:

Known at plan time

for_each keys and count values must be known when Terraform plans, because they decide how many instances exist. This fails on a fresh environment:

A network interface (NIC) is a VM's network card (1.13) - in Azure it is a separate resource attached to a subnet.

resource "azurerm_network_interface" "nic" {
  for_each = toset(azurerm_subnet.s[*].id)    # IDs do not exist yet
  # ...
}
│ Error: Invalid for_each argument
│
│ The "for_each" set includes values derived from resource attributes that
│ cannot be determined until apply, and so Terraform cannot determine the full
│ set of keys that will identify the instances of this resource.
│
│ When working with unknown values in for_each, it's better to use a map value
│ where the keys are defined statically in your configuration and where only
│ the values contain apply-time results.

The fix is exactly what the message says: key by something you wrote (the subnet name), and put the unknown ID in the value:

for_each  = azurerm_subnet.s            # a map keyed by subnet name
subnet_id = each.value.id               # the unknown part is only in the value

A for_each resource is a map, so you can chain one directly off another.

Nested collections: the flatten pattern

Two VNets, each with its own subnets. (A common design called hub and spoke: one central "hub" network with shared things like a firewall, and "spoke" networks for applications, connected to it.) for_each needs one flat map with one entry per subnet:

variable "vnets" {
  default = {
    hub   = { cidr = "10.0.0.0/16", subnets = { firewall = "10.0.1.0/24", bastion = "10.0.2.0/26" } }
    spoke = { cidr = "10.1.0.0/16", subnets = { app = "10.1.1.0/24", data = "10.1.2.0/24" } }
  }
}

locals {
  subnets = flatten([
    for vnet_name, vnet in var.vnets : [
      for sub_name, cidr in vnet.subnets : {
        key  = "${vnet_name}.${sub_name}"
        vnet = vnet_name
        name = sub_name
        cidr = cidr
      }
    ]
  ])
}

resource "azurerm_virtual_network" "v" {
  for_each            = var.vnets
  name                = "vnet-${each.key}"
  address_space       = [each.value.cidr]
  location            = var.location
  resource_group_name = azurerm_resource_group.main.name
}

resource "azurerm_subnet" "s" {
  for_each             = { for s in local.subnets : s.key => s }
  name                 = each.value.name
  virtual_network_name = azurerm_virtual_network.v[each.value.vnet].name
  address_prefixes     = [each.value.cidr]
  resource_group_name  = azurerm_resource_group.main.name
}
azurerm_subnet.s["hub.bastion"]
azurerm_subnet.s["hub.firewall"]
azurerm_subnet.s["spoke.app"]
azurerm_subnet.s["spoke.data"]

The inner for builds a list of objects per VNet; the outer builds a list of lists; flatten makes it one list; the final for keys it. Remove spoke.app and exactly one subnet is destroyed.

Later (Ch 13): modules (reusable groups of resources) take count and for_each too, with the same addressing: module.env["prod"].azurerm_resource_group.this.

Inside the instance

count.index      0, 1, 2 ...           only with count
each.key         the map key / set element   only with for_each
each.value       the map value / the same element for a set

Referencing instances from outside:

azurerm_subnet.s[0].id                        # count: by index
azurerm_subnet.s["web"].id                    # for_each: by key
[for s in azurerm_subnet.s : s.id]            # all of a for_each resource
azurerm_subnet.s[*].id                        # all of a count resource
values(azurerm_subnet.s)[*].id                # all of a for_each resource, as a list
{ for k, s in azurerm_subnet.s : k => s.id }  # keep the keys

If you already used count

State (12.1) records each instance under its address. terraform state mv FROM TO renames an entry in state - the real object is untouched, so nothing is destroyed (Ch 13 covers it fully):

terraform state mv 'azurerm_subnet.s[0]' 'azurerm_subnet.s["web"]'
terraform state mv 'azurerm_subnet.s[1]' 'azurerm_subnet.s["app"]'

or, reviewably in a PR, moved blocks - the same rename written in the code, applied by the next apply:

moved {
  from = azurerm_subnet.s[0]
  to   = azurerm_subnet.s["web"]
}

Either way the check is the same: the next plan shows moves and no replacements. Tedious for five subnets, and worth it for anything you cannot afford to recreate.

A review checklist for count and for_each

What you can now do:

Why it helps

This is the single most common Terraform outage pattern, and it arrives as an innocent PR: "remove the unused app subnet from the list". The plan says Plan: 3 to add, 0 to change, 4 to destroy, and if nobody reads it, four subnets and whatever is in them go down. Knowing this lets you block it in review and propose the for_each fix plus moved blocks to migrate without recreating. You will also hit "Invalid for_each argument" for lists and for keys known only after apply; knowing the "keys static, values unknown" rule fixes both. Interviewers ask count vs for_each constantly because it separates people who have run Terraform in production from those who have not.

FAQ

Is count ever the right choice?

Yes, for two cases. The on/off switch: count = var.enable_bastion ? 1 : 0, read with one(resource[*].attr). And genuinely identical, anonymous copies where you never remove one from the middle. Even then many teams use for_each keyed by a name, because "identical workers" eventually stop being identical. Anything with an identity (a name, a CIDR, a purpose) should use for_each.

Why does for_each reject my list?

for_each needs a map or a set of strings, because each instance needs a unique key and a list allows duplicates and uses positions. Convert with toset(var.names) for strings, where each.key and each.value are the same. For a list of objects, build a map: { for s in var.subnets : s.name => s }, choosing a key that is unique, stable and known at plan time.

What does "cannot be determined until apply" mean for for_each?

The keys decide how many instances exist, so Terraform needs them at plan time. If the keys come from IDs of resources not yet created, it cannot plan. The fix is to key by something you wrote (a name) and keep the unknown ID in the value. A for_each resource is itself a map keyed by your static keys, so for_each = azurerm_subnet.s with each.value.id works.

I already used count. How do I switch without destroying everything?

Re-address the existing instances in state. In a PR, write moved { from = azurerm_subnet.s[0] to = azurerm_subnet.s["web"] } for each one, or run terraform state mv with the same addresses (quote them). Then plan: the goal is moves only and no replacements. moved blocks are better for shared code because they are reviewed in the PR and apply everywhere the code is used.

How do I use for_each with nested data like VNets and their subnets?

Flatten it. A for inside a for builds a list of objects per VNet, flatten makes it one list, and a final { for s in local.subnets : s.key => s } keys it with something like "hub.firewall". Each subnet then has its own stable address, and removing one destroys exactly one. The same pattern works for any parent-child structure: permissions per team, DNS records per zone.

In an interview Junior

What is the difference between count and for_each, and why does it matter?

Both create several instances of a resource; the difference is how each instance is addressed, and the address is the identity Terraform matches to state:

Remove the second item of a list under count and every element after it shifts down an index; anything whose name cannot change in place is replaced. One removed subnet becomes four destroyed. With for_each, removing "app" is exactly one destroy.

Rule: for_each for anything with an identity, keyed by something unique, stable and known at plan time (a name, not an ID that exists only after apply). count for identical anonymous copies and the on/off switch (count = var.enabled ? 1 : 0, read with one(...[*])).

Already on count? Rename the state entries with moved blocks, and check the plan shows no replacements.

Also asked: Why does for_each not accept a list? · What does "known at plan time" mean for for_each keys? · How do you migrate a resource from count to for_each without recreating it?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.