OnCallReady

Lesson 30.14 · AWS II: VPC, EC2, ELB & EKS · 14 min read

EC2: images, instance types, user data, EBS

In plain words

An EC2 instance is a rented computer you order from a form. The form asks: which disk image to copy (the AMI), how big a machine (the instance type), which room to put it in (subnet), which door rules (security groups), which badge it wears (the instance profile), and a note with first-day instructions (user data) that it reads exactly once when it first wakes up. Its hard disk (EBS) is a network disk you can make bigger while it runs.

EC2: images, instance types, user data, EBS

An EC2 instance is a virtual machine you rent by the second. What makes running them sane is that everything about one is decided at launch, in a handful of parameters you can read back later. This lesson goes through those parameters, the instance's life cycle, and its disk.

Need to know: an instance = AMI (the disk image: OS + whatever was baked in) + instance type (CPU, memory, architecture) + subnet + security groups + instance profile (its IAM role) + user data (a script cloud-init runs once at first boot) + EBS volumes. Look AMIs up by owner and name, never by a pasted ID. The type's architecture must match the AMI's (Graviton t4g/m8g = arm64). EBS volumes grow online (modify-volume), but the partition and the file system have to be grown inside the instance too.

Finding the image

AMI IDs differ per Region and change with every build. Three ways to find the right one:

$ aws ec2 describe-images --owners self --filters 'Name=name,Values=oncall-web-*' --query 'sort_by(Images, &CreationDate)[].[ImageId,Name,CreationDate,Architecture]' --output table
----------------------------------------------------------------------------------------
|                                    DescribeImages                                    |
+-----------------------+------------------------+----------------------------+--------+
|  ami-09a93c85f343a8fa3|  oncall-web-2026.09.3  |  2026-09-17T14:40:52.000Z  |  arm64 |
|  ami-068ae5a871936e07f|  oncall-web-2026.10.1  |  2026-10-01T09:02:11.000Z  |  arm64 |
+-----------------------+------------------------+----------------------------+--------+
$ aws ec2 describe-images --owners 099720109477 --filters 'Name=name,Values=ubuntu/images/hvm-ssd-gp3/ubuntu-resolute-26.04-arm64-server-*' --query 'sort_by(Images, &CreationDate)[-1].[ImageId,Name]' --output text
ami-04fa9a8d22329866c	ubuntu/images/hvm-ssd-gp3/ubuntu-resolute-26.04-arm64-server-20260924
$ aws ssm get-parameter --name /aws/service/canonical/ubuntu/server/26.04/stable/current/arm64/hvm/ebs-gp3/ami-id --query Parameter.Value --output text
ami-04fa9a8d22329866c
  1. Your own images (--owners self): the team's "golden" images with the app baked in. sort_by(..., &CreationDate)[-1] is the newest - the list order is not age order.
  2. A publisher's images by owner account (099720109477 = Canonical, amazon = Amazon Linux) and a name pattern.
  3. SSM public parameters: Canonical and AWS publish the current image ID under a fixed name, the simplest way for scripts and Terraform (data "aws_ssm_parameter").

Never take an AMI from an unknown owner: anyone can publish a public AMI named "ubuntu-...".

Instance types

t4g.small = family t (burstable), generation 4, g = Graviton (AWS's arm64 CPU), size small. Families: t burstable (CPU credits - fine for spiky, wrong for steady load), m general purpose, c compute, r memory; g arm64, a AMD, no letter Intel (m7i).

$ aws ec2 describe-instance-types --instance-types t4g.small m7g.large c7i.large --query 'InstanceTypes[].[InstanceType,VCpuInfo.DefaultVCpus,MemoryInfo.SizeInMiB,ProcessorInfo.SupportedArchitectures[0]]' --output table
--------------------------------------
|        DescribeInstanceTypes       |
+------------+----+-------+----------+
|  t4g.small |  2 |  2048 |  arm64   |
|  m7g.large |  2 |  8192 |  arm64   |
|  c7i.large |  2 |  4096 |  x86_64  |
+------------+----+-------+----------+

Graviton is ~20% cheaper per hour for the same size and usually faster; the only catch is that the AMI and every binary must be arm64. An x86 AMI on a t4g fails at launch:

$ SUB=$(aws ec2 describe-subnets --filters Name=tag:Name,Values=try-private-a --query 'Subnets[0].SubnetId' --output text)
$ aws ec2 run-instances --image-id $(aws ec2 describe-images --owners 099720109477 --filters 'Name=name,Values=*resolute-26.04-amd64-server-*' --query 'Images[0].ImageId' --output text) --instance-type t4g.small --subnet-id $SUB
aws: [ERROR]: An error occurred (InvalidParameterValue) when calling the RunInstances operation: The architecture 'arm64' of the specified instance type does not match the architecture 'x86_64' of the specified AMI. Specify an instance type and an AMI that have matching architectures, and try again. You can use 'describe-instance-types' or 'describe-images' to discover the architecture of the instance type or AMI.

Reading an instance back

The lesson's instance try-web was launched 25 minutes ago with the team's image, no key pair, the oncall-web-ec2 profile and a small user data script:

$ aws ec2 describe-instances --filters Name=tag:Name,Values=try-web --query 'Reservations[].Instances[].{id:InstanceId,state:State.Name,type:InstanceType,az:Placement.AvailabilityZone,ip:PrivateIpAddress,public:PublicIpAddress,key:KeyName,profile:IamInstanceProfile.Arn,tokens:MetadataOptions.HttpTokens}' --output yaml
- id: i-0961d651f05141eaf
  state: running
  type: t4g.small
  az: eu-central-1a
  ip: 10.99.10.30
  public: null
  key: null
  profile: arn:aws:iam::111122223333:instance-profile/oncall-web-ec2
  tokens: required

public: null (a private subnet), key: null (no SSH key at all), tokens: required (IMDSv2 only - next lesson). The states: pending -> running -> stopping -> stopped (the EBS volumes stay, you pay for them, not for the CPU; a public IP is released) and shutting-down -> terminated (gone; its root volume is deleted with it by default). Stop is "turn it off", terminate is "delete it". Termination protection (--disable-api-termination) guards the ones that must not go.

User data and cloud-init

User data is a script (or a #cloud-config document) passed at launch. cloud-init runs it once, as root, at the first boot, late in the boot. Its output goes to /var/log/cloud-init-output.log inside, and to the serial console - which you can read from outside, without any access to the instance:

$ I=$(aws ec2 describe-instances --filters Name=tag:Name,Values=try-web --query 'Reservations[0].Instances[0].InstanceId' --output text)
$ aws ec2 describe-instance-attribute --instance-id $I --attribute userData --query UserData.Value --output text | base64 -d
#!/bin/bash
set -euo pipefail
apt-get update
apt-get install -y jq
echo "try-web: user data done"
$ aws ec2 get-console-output --instance-id $I --latest --output text | grep -E 'cloud-init|user data|Setting up' | tail -6
[   14.157700] cloud-init[1044]: The following NEW packages will be installed:
[   14.576000] cloud-init[1044]:   jq
[   14.994300] cloud-init[1044]: 0 upgraded, 1 newly installed, 0 to remove and 0 not upgraded.
[   15.412600] cloud-init[1044]: Setting up jq (1.8.1-3) ...
[   15.830900] cloud-init[1044]: try-web: user data done
[   16.249200] cloud-init[1044]: Cloud-init v. 25.1.4-0ubuntu1~26.04.1 finished at Tue, 22 Sep 2026 19:35:40 +0000. Datasource DataSourceEc2Local.  Up 33.50 seconds

Two consequences people learn the hard way: changing the user data of an existing instance does nothing (it already ran); and a script that installs from the internet fails silently on an instance without a way out (a lab in this chapter). Golden images move that work to build time, so boot does almost nothing and cannot fail on a mirror.

Launch templates store all launch parameters as a versioned object (create-launch-template, create-launch-template-version); Auto Scaling groups require one. run-instances --launch-template LaunchTemplateName=web,Version=3 launches from it by hand.

EBS volumes

The root disk is an EBS volume: network storage in the instance's AZ (it cannot attach to an instance in another AZ), replicated inside that AZ, snapshot to S3.

$ aws ec2 describe-volumes --filters Name=attachment.instance-id,Values=$I --query 'Volumes[].[VolumeId,Size,VolumeType,Iops,Throughput,Attachments[0].Device,Attachments[0].DeleteOnTermination]' --output table
---------------------------------------------------------------------------
|                             DescribeVolumes                             |
+------------------------+----+------+-------+------+------------+--------+
|  vol-082a42b2efc9a5974 |  8 |  gp3 |  3000 |  125 |  /dev/sda1 |  True  |
+------------------------+----+------+-------+------+------------+--------+

gp3 is the default and the one to use: 3,000 IOPS and 125 MB/s baseline at any size, more of either bought separately (unlike gp2, whose IOPS grew with size). io2 for databases that need guaranteed IOPS; st1/sc1 HDDs for throughput-heavy cold data.

Growing a full disk is three layers, from outside in:

LayerCommandWhere
volumeaws ec2 modify-volume --volume-id vol-... --size 16 (online; up only; once per 6 h)your machine
partitionsudo growpart /dev/nvme0n1 1 (disk, space, partition number)the instance
file systemsudo resize2fs /dev/nvme0n1p1 (ext4) / sudo xfs_growfs / (XFS)the instance

On Nitro instances EBS volumes show up as NVMe devices (/dev/nvme0n1), whatever the "device name" /dev/sda1 in the API says. lsblk inside shows the three layers side by side.

In an interview: "The root disk of an EC2 instance is full - how do you grow it without downtime?" - modify-volume to the new size (online), wait until the modification is optimizing, then growpart the partition and resize2fs (or xfs_growfs) the file system, and check df -h.

What you can do now

Why it helps

Every instance-related incident traces back to one of the launch parameters: the wrong image or architecture, a subnet without a way out, a security group, a missing role, a user data script that failed at first boot, or a full disk. Reading those parameters back, debugging user data from outside with the console output, and growing a disk through all three layers are skills you use on call every month.

Commands in this lesson

aws

FAQ

How do I pick the right AMI?

Look it up instead of pasting it: your team's golden images with --owners self and a name filter, a publisher's images by their owner account (Canonical is 099720109477), or the SSM public parameter for the current Ubuntu or Amazon Linux image. Sort by CreationDate for the newest. Never trust an AMI from an unknown owner just because of its name.

What is the difference between stop and terminate?

Stop shuts the instance down; its EBS volumes stay (and you keep paying for them), its auto-assigned public IP is released, and it can be started again. Terminate deletes it, and by default its root volume with it. Termination protection guards instances that must never be deleted by mistake.

What is Graviton and why does the AMI matter?

Graviton is AWS's own arm64 CPU (instance types with a g, like t4g or m7g). It is usually about 20% cheaper for the same size. The AMI and every binary must be built for arm64: launching an x86 image on a t4g fails with an architecture mismatch error.

Why did changing the user data do nothing?

User data runs once, at the first boot, through cloud-init. Changing it on an existing instance does not run it again. To change what a fleet does at boot, change the launch template and replace the instances (an instance refresh); better still, bake the setup into the image so boot does almost nothing.

Why does df still show 8G after I resized the volume?

Because three layers need to grow: the EBS volume (modify-volume), the partition on it (growpart), and the file system in the partition (resize2fs for ext4, xfs_growfs for XFS). lsblk inside the instance shows the disk and partition sizes side by side, which tells you which step is missing.

In an interview Mid

The root disk of an EC2 instance is full. How do you grow it without downtime?

Grow the EBS volume with aws ec2 modify-volume --size while the instance runs (only up, once per six hours; usable once the modification is optimizing). Then inside the instance grow the partition with growpart (disk, space, partition number) and the file system with resize2fs for ext4 or xfs_growfs for XFS, and confirm with df -h. lsblk shows which layer is still small. Then ask why the data is on the root volume.

Also asked: Where do you look when an instance did not do what its user data says? · What is the difference between gp2 and gp3 volumes? · Why use a launch template instead of run-instances parameters?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.