OnCallReady

Lesson 30.12 · AWS II: VPC, EC2, ELB & EKS · 16 min read

AWS with Terraform: the aws provider and state in S3

In plain words

Terraform is a recipe book: you write down what the kitchen should contain, and the cook (the aws provider) makes the API calls to get there, writing in a notebook (the state) what it made. On AWS the notebook lives in a shared S3 bucket so everyone uses the same one, and whoever is cooking puts a "busy" sign (the lock file) on it so two cooks never write at once.

AWS with Terraform: the aws provider and state in S3

You learned Terraform on Azure. On AWS the language, the workflow and the state are the same; what changes is the provider (hashicorp/aws, 6.x), a handful of AWS-specific habits, and where the team keeps its state. This lesson runs a small network project and then looks at the parts that differ.

Need to know: configure the provider with a region (credentials come from the same chain as the CLI: AWS_PROFILE, environment keys, the instance role). Resource IDs (vpc-..., subnet-...) only exist after the API call, so they are (known after apply) in a first plan. The team's state goes in S3 with backend "s3", a versioned bucket, and use_lockfile = true (Terraform 1.10+): the lock is a .tflock object next to the state; the DynamoDB lock table is deprecated.

A small project

~/oncall-lab/labs/aws/try/tf-net/main.tf has a VPC, two subnets spread over the AZs a data source returns, and a security group:

$ cd ~/oncall-lab/labs/aws/try/tf-net
$ grep -n -A3 'provider "aws"\|data "\|resource "' main.tf | head -40
10:provider "aws" {
11-  region = "eu-central-1"
12-  default_tags {
13-    tags = { managed-by = "terraform", lab = "try" }
--
17:data "aws_caller_identity" "me" {}
18:data "aws_availability_zones" "available" {
19-  state = "available"
20-}
21-
22:resource "aws_vpc" "try" {
23-  cidr_block = "10.98.0.0/16"
24-  tags       = { Name = "try-tf-vpc" }
25-}
--
27:resource "aws_subnet" "app" {
28-  count             = 2
29-  vpc_id            = aws_vpc.try.id
30-  cidr_block        = cidrsubnet(aws_vpc.try.cidr_block, 8, count.index)
--
35:resource "aws_security_group" "app" {
36-  name   = "try-tf-app"
37-  vpc_id = aws_vpc.try.id
38-  ingress {
$ terraform init -no-color | grep -E 'Installing|Installed|initialized'
- Installing hashicorp/aws v6.68.0...
- Installed hashicorp/aws v6.68.0 (signed by HashiCorp)
Terraform has been successfully initialized!

data "aws_availability_zones" asks EC2 which AZs the Region has, so the code works in any Region; count + cidrsubnet(cidr, 8, count.index) turns a /16 into /24s; default_tags puts the same tags on every resource the provider creates (and they show up in tags_all, not in each tags).

$ terraform plan -no-color | grep -E '^  # |known after apply|Plan:' | head -20
  # aws_vpc.try will be created
      + arn                       = (known after apply)
      + default_network_acl_id    = (known after apply)
      + default_route_table_id    = (known after apply)
      + default_security_group_id = (known after apply)
      + dhcp_options_id           = (known after apply)
      + enable_dns_hostnames      = (known after apply)
      + enable_dns_support        = (known after apply)
      + id                        = (known after apply)
      + instance_tenancy          = (known after apply)
      + main_route_table_id       = (known after apply)
      + owner_id                  = (known after apply)
      + tags_all                  = (known after apply)
  # aws_subnet.app[0] will be created
      + arn                     = (known after apply)
      + availability_zone_id    = (known after apply)
      + id                      = (known after apply)
      + map_public_ip_on_launch = (known after apply)
      + owner_id                = (known after apply)
      + tags_all                = (known after apply)

The IDs and ARNs are (known after apply): AWS assigns them. Terraform creates in dependency order (VPC, then subnets and the group, which reference aws_vpc.try.id) and passes the real IDs along:

$ terraform apply -auto-approve -no-color | grep -E 'Creation complete|Apply complete|account|subnet-'
  + account = "111122223333"
aws_vpc.try: Creation complete after 4s [id=vpc-0fab6173e91281414]
aws_subnet.app[0]: Creation complete after 4s [id=subnet-07d93f499311787f9]
aws_subnet.app[1]: Creation complete after 4s [id=subnet-013c96da193a52f19]
aws_security_group.app: Creation complete after 4s [id=sg-026e0a574173c3373]
Apply complete! Resources: 4 added, 0 changed, 0 destroyed.
account = "111122223333"
          "subnet-07d93f499311787f9",
          "subnet-013c96da193a52f19",
$ aws ec2 describe-subnets --filters Name=tag:Name,Values='try-tf-app-*' --query 'Subnets[].[Tags[?Key==`Name`]|[0].Value,CidrBlock,AvailabilityZone]' --output text
try-tf-app-0	10.98.0.0/24	eu-central-1a
try-tf-app-1	10.98.1.0/24	eu-central-1b

The CLI sees exactly what Terraform made - the provider makes the same API calls, as the same identity.

Habits that differ from azurerm

The egress rule. AWS gives every new security group an allow-all outbound rule. The Terraform provider removes it: an aws_security_group without egress blocks allows nothing out. That is deliberate (the code should say what is allowed) and it is the first "my instance cannot reach anything" bug for people coming from the console:

$ aws ec2 describe-security-groups --filters Name=group-name,Values=try-tf-app --query 'SecurityGroups[0].IpPermissionsEgress'
[]

Inline rules vs rule resources. Rules can be written inline (ingress {} blocks, the group owns the whole list) or as separate aws_vpc_security_group_ingress_rule resources (one rule each, the recommended style in 6.x). Never mix both for the same group: each side removes the other's rules on every apply.

Drift is real API state. Change something with the CLI and the next plan shows it - and with inline rules, plans to take it back out:

$ SG=$(aws ec2 describe-security-groups --filters Name=group-name,Values=try-tf-app --query 'SecurityGroups[0].GroupId' --output text)
$ aws ec2 authorize-security-group-ingress --group-id $SG --protocol tcp --port 22 --cidr 0.0.0.0/0 --query Return
true
$ terraform plan -no-color | grep -E 'has changed|will be updated|from_port|Plan:'
  # aws_security_group.app has changed
          "from_port"   = 8080
            "from_port"        = 8080
            "from_port"        = 22
  # aws_security_group.app will be updated in-place
            "from_port"        = 8080
            "from_port"        = 22
          "from_port"   = 8080
Plan: 0 to add, 1 to change, 0 to destroy.

That plan is the reason to keep production changes in code: the hand-added SSH rule would be reverted by the next apply, and the review shows it.

Look-ups instead of hard-coded IDs. data "aws_ami" with owners and a name filter (most_recent = true), data "aws_vpc" by tag, data "aws_caller_identity" for the account ID in ARNs. AMI IDs differ per Region and change with every image build; never paste one into code.

State in S3

terraform {
  backend "s3" {
    bucket       = "oncall-tfstate-111122223333"
    key          = "network/terraform.tfstate"
    region       = "eu-central-1"
    use_lockfile = true
  }
}
$ aws s3api get-bucket-versioning --bucket oncall-tfstate-111122223333
{
    "Status": "Enabled"
}

A lock left behind by a run that died (a killed CI job) is removed with terraform force-unlock <LOCK_ID> - after checking that the run really is gone.

In an interview: "How do you store Terraform state for a team on AWS?" - an S3 backend in a versioned, private, encrypted bucket, one key per stack, with locking (use_lockfile = true on Terraform 1.10+, a DynamoDB table before that), and CI as the only writer.

$ terraform destroy -auto-approve -no-color | tail -1
Destroy complete! Resources: 4 destroyed.

What you can do now

Why it helps

Most AWS teams manage their networks and clusters with Terraform, and most Terraform incidents on AWS are not about the language: they are a security group without egress, inline rules fighting separate rule resources, hard-coded AMI IDs, or two applies corrupting a state with no lock. Knowing the aws provider's habits and the S3 backend with use_lockfile makes your code safe to run by a team and by CI.

Commands in this lesson

cd grep terraform aws

FAQ

Why are so many values "known after apply"?

Because AWS assigns them: VPC, subnet and instance IDs, ARNs and addresses only exist after the API call returns. Terraform plans with placeholders, creates resources in dependency order and passes the real IDs to the resources that reference them. A second plan after the apply shows the real values and, if nothing drifted, no changes.

Why does my security group created by Terraform block all outbound traffic?

AWS gives every new security group an allow-all outbound rule, but the Terraform aws provider removes it so that the code describes everything that is allowed. A group without egress blocks or egress rule resources therefore allows nothing out. Add the egress you need explicitly.

What does use_lockfile do?

With use_lockfile = true (Terraform 1.10 and later) the S3 backend writes a .tflock object next to the state with a conditional PUT at the start of every run. A second run gets PreconditionFailed and stops with the lock's owner and ID instead of overwriting the state. The run deletes the object when it finishes. It replaces the DynamoDB lock table, which still works but is deprecated.

Why version the state bucket?

Every apply writes a new state; with versioning, the previous ones stay as older object versions. If an apply goes wrong or someone edits the state by hand, you can restore an earlier version. Combine it with Block Public Access, encryption and a policy that only lets CI and a few admins write.

How do I avoid hard-coding AMI IDs?

Use a data source: data "aws_ami" with owners, a name filter and most_recent = true, or data "aws_ssm_parameter" for the public parameters Canonical and AWS publish. AMI IDs differ per Region and change with every image build, so a pasted ID is wrong in another Region and outdated next month.

In an interview Mid

How do you store Terraform state for a team on AWS?

In an S3 backend: a dedicated bucket with versioning, Block Public Access and encryption, one key per stack and environment, and locking with use_lockfile = true on Terraform 1.10+ (a .tflock object written with a conditional PUT; a DynamoDB table on older versions). CI is the only regular writer, people use plan locally, and terraform force-unlock is only for locks whose run is really gone.

Also asked: What happens when someone changes a Terraform-managed security group in the console? · Why should you not mix inline security group rules with separate rule resources? · How do you pass outputs from a network stack to an application stack?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.