Why this matters
What you need to know already: resource blocks and references (12.1), outputs (12.5), expressions (12.11), dependencies between systemd units (2.14 - "ordering is not dependency" has a Terraform twin here).
Your configuration rarely lives alone. The shared network was built by another team; the account you run as already exists. You need their IDs without taking ownership - and Terraform must create your own things in the right order. This lesson covers both: data sources (read-only lookups) and the dependency graph.
resource vs data, precisely
resource "azurerm_resource_group" "orders" { # Terraform OWNS this
name = "rg-orders-dev"
location = "westeurope"
}
data "azurerm_resource_group" "shared" { # Terraform READS this
name = "rg-shared-network"
}
A resource is created, updated and destroyed by this configuration. It is in state as an object Terraform manages, and it appears in destroy plans.
A data source is a query. Terraform asks the provider "what does this look like right now?" on every plan, uses the answer, and forgets it. It never creates, changes or deletes the target. terraform destroy shows only resources.
The two share a type name - azurerm_resource_group exists as both - but the data source takes fewer arguments (the ones that identify the object) and exposes the rest as attributes:
output "shared_location" {
value = data.azurerm_resource_group.shared.location
}
Note the address: data sources are always prefixed with data..
When to use a data source
- Something another team owns. The central ("hub") network, the shared logging workspace, the DNS zone. You need its ID; you must not manage it.
- Something that exists before Terraform. The subscription, the tenant (the company's Azure directory of users and robot accounts), the identity running Terraform.
- Facts about the platform. The latest VM image, the available versions of a service, who you are logged in as.
The one every Azure configuration has:
data "azurerm_client_config" "current" {}
resource "azurerm_key_vault" "main" {
name = "kv-orders-dev"
location = azurerm_resource_group.orders.location
resource_group_name = azurerm_resource_group.orders.name
tenant_id = data.azurerm_client_config.current.tenant_id
sku_name = "standard"
}
azurerm_client_config needs no arguments; it returns tenant_id, subscription_id, client_id and object_id of whoever is running Terraform - which is how you grant the CI job's own identity access to the Key Vault (secret store, 12.3) it just created. sku_name is the vault's price tier.
Looking up something that does not exist is a plan-time error, which is exactly what you want - "the hub VNet is not where we think it is" should stop the run, not create a new one.
data vs import, the question people get wrong
If you want Terraform to take over an existing resource - manage it from now on, change it, eventually delete it - that is import, not a data source. Importing gives you a resource block you own. A data source gives you read-only attributes of something you will never touch. Ch 13 does import.
When data sources are read
Normally during plan, before anything is created. (azurerm_subnet below looks up a subnet - an address range inside a VNet, 12.8.) But if a data source's arguments depend on something that does not exist yet, it cannot be read until apply:
data "azurerm_subnet" "app" {
name = "snet-app"
virtual_network_name = azurerm_virtual_network.main.name # being created in this run
resource_group_name = azurerm_resource_group.orders.name
}
The plan shows it being deferred:
# data.azurerm_subnet.app will be read during apply
# (config refers to values not yet known)
<= data "azurerm_subnet" "app" {
+ id = (known after apply)
...
}
How to read this plan fragment: # lines are Terraform's comment on what will happen; <= is the "read" symbol (you will meet + create, ~ update, - destroy in 12.24); (known after apply) means the value does not exist yet. Everything computed from it is (known after apply) too. Reading a data source about something you create in the same configuration is usually a smell: reference the resource directly (azurerm_subnet.app.id) instead.
References build the graph
Terraform never runs your blocks in file order. It builds a dependency graph
- a map of "what needs what" - from references and walks it, doing independent
things in parallel (10 at a time by default; the flag -parallelism=N on plan or apply changes it).
resource "azurerm_resource_group" "main" {
name = "rg-orders-dev"
location = "westeurope"
}
resource "azurerm_virtual_network" "main" {
name = "vnet-orders-dev"
location = azurerm_resource_group.main.location # a reference
resource_group_name = azurerm_resource_group.main.name # another one
address_space = ["10.20.0.0/16"]
}
resource "azurerm_storage_account" "logs" {
name = "stordersdevlogs"
resource_group_name = "rg-orders-dev" # a string, not a reference!
location = "westeurope"
account_tier = "Standard"
account_replication_type = "LRS"
}
The VNet implicitly depends on the resource group because it references it. The storage account does not - it names the resource group with a literal string, so Terraform may try to create it at the same time as the group, and on a fresh environment it fails with ResourceGroupNotFound. The fix is not depends_on; it is referencing the resource (azurerm_resource_group.main.name). References carry the dependency and the value at once.
You can see the graph. terraform graph prints it:
$ terraform graph
digraph G {
rankdir = "RL";
node [shape = rect, fontname = "sans-serif"];
"azurerm_resource_group.main" [label="azurerm_resource_group.main"];
"azurerm_storage_account.logs" [label="azurerm_storage_account.logs"];
"azurerm_virtual_network.main" [label="azurerm_virtual_network.main"];
"azurerm_virtual_network.main" -> "azurerm_resource_group.main";
}
The output is DOT, a text format for drawings: each "A" -> "B" line is an arrow "A depends on B". Paste it into any Graphviz viewer to see boxes and arrows. The storage account floating on its own, with no arrow, is the bug above made visible.
Destroy walks the same graph backwards: the VNet goes before the resource group, because it depends on it.
depends_on: only for dependencies Terraform cannot see
A role assignment is Azure's permission grant: "this identity may do this role's actions on this thing" - like adding a user to a group (4.3), for the cloud. Here, the identity running Terraform gets permission to write secrets into the vault:
resource "azurerm_role_assignment" "pipeline_kv" {
scope = azurerm_key_vault.main.id
role_definition_name = "Key Vault Secrets Officer"
principal_id = data.azurerm_client_config.current.object_id
}
resource "azurerm_key_vault_secret" "db" {
name = "db-password"
value = random_password.db.result
key_vault_id = azurerm_key_vault.main.id
depends_on = [azurerm_role_assignment.pipeline_kv]
}
The secret references the vault, but nothing in it references the role assignment - yet writing the secret fails with 403 (HTTP "forbidden", 9.21) until the role exists. That is a hidden dependency: it lives in Azure's authorisation model, not in any attribute. depends_on states it explicitly.
Rules for depends_on:
- It takes a list of references (no quotes): resources, data sources, or whole modules (Ch 13).
- Use it only for what Terraform genuinely cannot infer. It is not a way to control ordering for the sake of it.
- On a data source it forces the read to wait until apply - every plan then shows
(known after apply)for it. Avoid it there. - On a module it makes everything in the module wait for everything in the dependency, which makes large parts of the graph run one at a time.
Meta-arguments
Arguments that every resource accepts, because Terraform - not the provider - handles them:
depends_on explicit dependencies
count N instances, addressed [0], [1], ...
for_each one instance per map key / set element, addressed ["key"]
provider which provider configuration: provider = azurerm.dr
lifecycle create_before_destroy, prevent_destroy, ignore_changes,
replace_triggered_by, precondition, postcondition
(count and for_each are 12.20; lifecycle is Ch 13.) Modules accept count, for_each, depends_on and providers. That is the complete list; anything else in a block is an argument for the provider.
Addresses, precisely
An address names one thing in your configuration. You will type addresses into terraform state, -target, -replace, import and moved blocks constantly (12.24 and Ch 13):
azurerm_subnet.app a resource (all instances)
azurerm_subnet.app[0] one instance (count)
azurerm_subnet.app["web"] one instance (for_each)
data.azurerm_client_config.current a data source
module.network a module call
module.network.azurerm_subnet.app["web"] a resource inside it
module.env["prod"].module.network.azurerm_subnet.app["web"] nested, with for_each
In a shell, quote addresses that contain brackets and double quotes:
terraform state show 'azurerm_subnet.app["web"]'
terraform apply -replace='azurerm_linux_virtual_machine.web[1]'
(terraform state show ADDR prints what state holds for one resource; -replace=ADDR forces one resource to be rebuilt. Both come back in 12.24.) Without the single quotes, bash eats the double quotes (6.6) and zsh tries to glob the brackets.
What you can now do:
- Choose between a
resource(own it) and adatasource (read it). - Explain why references, not file order, decide what is created first.
- Use
depends_ononly for hidden dependencies, and quote addresses in the shell.