The problem: Vault configured by hand
Every policy, mount, auth method and role you made in this chapter was a command typed in a terminal. A year later nobody knows why ci-deploy can read secret/data/shop/*, the dev and prod Vaults differ in ways nobody chose, and a policy deleted by mistake at 11 pm is gone until someone remembers what it said. The same answer as for cloud resources applies: Vault's configuration is code, reviewed in pull requests and applied by Terraform. And the same trap as in the Terraform chapters waits: state holds secrets in plain text.
What you need to know already: providers, versions and the lock file (12.3), state and why it holds secrets (13.1), backends and who may read them (13.4), drift and -refresh-only (13.21), secrets in Terraform with Key Vault (14.20), and this chapter's policy, auth method and KV v2 lessons.
The words you need first
- Configuration vs secret values - mounts, policies, auth methods and roles are configuration: they say who may do what. The passwords and keys inside KV are values. Terraform is the right tool for the first; for the second, it is a compromise.
- Child token - a token created by another token, revoked with it. The Vault provider does not run with your token directly: by default it creates a short-lived child of it for the run.
- Ephemeral resource (Terraform 1.10+) - a block that reads a value during plan and apply and never writes it to state or the plan file.
- Write-only argument (Terraform 1.11+) - an argument a provider accepts but never stores in state; you bump a companion version argument to send a new value.
The provider
terraform {
required_providers {
vault = {
source = "hashicorp/vault"
version = "~> 5.12"
}
}
}
provider "vault" {
address = "https://127.0.0.1:8200"
}
- No token in the block: the provider reads
VAULT_TOKEN, then~/.vault-token- the same places the CLI reads. In a pipeline that token comes from the CI system's Vault login (JWT/OIDC or AppRole), never from a variable in the repository. - By default the provider creates a child token of yours with a 20-minute TTL (
max_lease_ttl_seconds = 1200) and uses that for the run, so a leaked plan log holds a token that is already dead, and the audit log showstoken-terraform.skip_child_token = trueturns it off (needed when your token may not create children, e.g. a batch token). (simulator) The lab's provider uses your token directly. terraform initinstalls it like any other provider:
$ terraform init
Initializing the backend...
Initializing provider plugins...
- Finding hashicorp/vault versions matching "~> 5.12"...
- Installing hashicorp/vault v5.12.0...
- Installed hashicorp/vault v5.12.0 (signed by HashiCorp)
Terraform has created a lock file .terraform.lock.hcl to record the provider
selections it made above. Include this file in your version control repository
so that Terraform can guarantee to make the same selections by default when
you run "terraform init" in the future.
Terraform has been successfully initialized!
The resources you will write
| resource | does what the CLI did with |
|---|---|
vault_mount | vault secrets enable -path=... TYPE |
vault_auth_backend | vault auth enable TYPE |
vault_policy | vault policy write NAME FILE |
vault_approle_auth_backend_role | vault write auth/approle/role/NAME ... |
vault_kubernetes_auth_backend_config / _role | vault write auth/kubernetes/config / role/NAME |
vault_database_secret_backend_connection / _role | vault write database/config/NAME / roles/NAME |
vault_kv_secret_v2 | vault kv put |
vault_generic_endpoint | vault write PATH for anything without its own resource |
A policy and an AppRole role for an ETL job:
resource "vault_policy" "reports_read" {
name = "reports-read"
policy = <<-EOT
path "secret/data/reports/*" {
capabilities = ["read"]
}
EOT
}
resource "vault_auth_backend" "approle" {
type = "approle"
}
resource "vault_approle_auth_backend_role" "reports_etl" {
backend = vault_auth_backend.approle.path
role_name = "reports-etl"
token_policies = [vault_policy.reports_read.name]
token_ttl = 1200
token_max_ttl = 3600
}
References (vault_auth_backend.approle.path, vault_policy.reports_read.name) give Terraform the order: the auth method before the role, the policy before the role that names it. TTLs are seconds here, where the CLI accepted 20m.
$ terraform apply -auto-approve
...
vault_auth_backend.approle: Creating...
vault_auth_backend.approle: Creation complete after 4s [id=approle]
vault_approle_auth_backend_role.reports_etl: Creating...
vault_approle_auth_backend_role.reports_etl: Creation complete after 4s [id=auth/approle/role/reports-etl]
Apply complete! Resources: 2 added, 0 changed, 0 destroyed.
$ vault read -field=token_policies auth/approle/role/reports-etl; echo
[reports-read]
For a policy, the CLI's vault policy read and the resource hold the same text - and a policy file reviewed in a pull request is exactly the "least privilege, written down" this chapter has been asking for.
Drift: someone changes Vault by hand
Terraform compares its state with Vault at every plan. Delete the policy with the CLI (an operator cleaning up at 11 pm) and the next plan notices and puts it back:
$ vault policy delete reports-read
Success! Deleted policy: reports-read
$ terraform plan
vault_policy.reports_read: Refreshing state... [id=reports-read]
vault_kv_secret_v2.reports_db: Refreshing state... [id=secret/data/reports/db]
Note: Objects have changed outside of Terraform
Terraform detected the following changes made outside of Terraform since the
last "terraform apply" which may have affected this plan:
# vault_policy.reports_read has been deleted
- resource "vault_policy" "reports_read" {
- id = "reports-read" -> null
name = "reports-read"
# (1 unchanged attribute hidden)
}
...
# vault_policy.reports_read will be created
+ resource "vault_policy" "reports_read" {
...
Plan: 1 to add, 0 to change, 0 to destroy.
That is the value of Vault-as-code during an incident: "what did this policy say?" is a file in git, and putting it back is an apply - not memory.
Secrets in state
vault_kv_secret_v2 writes a secret value:
variable "db_password" {
type = string
sensitive = true
}
resource "vault_kv_secret_v2" "reports_db" {
mount = "secret"
name = "reports/db"
data_json = jsonencode({
DB_USER = "reports"
DB_PASSWORD = var.db_password
})
}
sensitive = true keeps the value out of the plan output, and the value comes from TF_VAR_db_password, not from a file in git. The plan looks clean:
# vault_kv_secret_v2.reports_db will be created
+ resource "vault_kv_secret_v2" "reports_db" {
+ data = (sensitive value)
+ data_json = (sensitive value)
+ id = (known after apply)
+ mount = "secret"
+ name = "reports/db"
+ path = (known after apply)
}
The state is not:
Everything you learned about state applies, with higher stakes: the backend must be encrypted and access-controlled (13.4), and anyone who can run terraform state pull can read every secret Terraform ever wrote. A data "vault_kv_secret_v2" block (reading a secret to pass it on) has the same effect: the value lands in state.
The modern answers
- Ephemeral resources (Terraform 1.10+, Vault provider 4.6+): read a secret without storing it.
ephemeral "vault_kv_secret_v2" "db" {
mount = "secret"
name = "reports/db"
}
# usable in provider blocks and write-only arguments, never stored
- Write-only arguments (Terraform 1.11+):
vault_kv_secret_v2takesdata_json_wo(never stored) plusdata_json_wo_version(bump it to send a new value). The secret is set in Vault and the state only has the version number.
resource "vault_kv_secret_v2" "reports_db" {
mount = "secret"
name = "reports/db"
data_json_wo = jsonencode({ DB_USER = "reports", DB_PASSWORD = var.db_password })
data_json_wo_version = 2
}
(simulator) The lab's Terraform implements neither ephemeral blocks nor write-only arguments; the code above is what you write on a real Terraform 1.11+ with provider 5.x.
- Do not let Terraform hold the value at all. The pattern most teams settle on: Terraform manages the structure (mounts, auth methods, policies, roles, database connections - with the root password rotated by Vault right after:
vault write -f database/rotate-root/NAME), and values come from elsewhere: dynamic secrets that no human sees, or a person / a rotation job writing KV values.
Two Vaults, one module
Dev and prod Vaults should differ only where you decided they differ: one module with the policies and roles, called per environment (14.1) with its own backend and its own provider token. A policy change is then reviewed once and applied to dev first.
In an interview: "I manage Vault's configuration - mounts, auth methods, policies, roles - with Terraform, but I keep secret values out of state: dynamic secrets, or ephemeral resources and write-only arguments on Terraform 1.11+, because state holds everything in plain text whatever sensitive says."
What you can now do
- Configure the
hashicorp/vaultprovider without a token in the code, and explain its child token. - Write
vault_policy,vault_auth_backendand an AppRole role as code, with references for ordering. - Show that
sensitive = truehides a value from the plan but not from the state, and find the value withjq. - Name the fixes: ephemeral resources, write-only arguments, and keeping values out of Terraform altogether.
- Use drift detection to restore a policy deleted by hand.