OnCallReady

Lesson 31.29 · AWS III: CloudWatch, CloudTrail, Cost & Incidents · 18 min read

EBS from the instance: gp3, growing a disk online, and the disk metrics EC2 does not have

In plain words

An EBS volume is a hard disk that AWS lends to your machine over the network. Unlike a real disk, you can make it bigger while the machine keeps running.

But the machine's own map of the disk does not change by itself. After AWS makes the disk bigger, you tell the machine twice: first stretch the partition (the slice of the disk) to the new end, then stretch the filesystem that lives inside the partition. No restart needed.

EBS from the instance: growing a disk online, and the metrics EC2 does not have

"Disk full" is the oldest page in operations, and on EC2 it has a good answer: an EBS volume can be made bigger while the instance runs, and the partition and filesystem can follow without a reboot. This lesson does it on ingest-1 - a managed host in the lab that plays the EC2 instance i-0d4e5f60718293a41 (simulator: AWS II brings real instances; the disk mechanics are the same) - and shows why CloudWatch alone never told you the disk was filling up.

Need to know: a gp3 volume has its size, IOPS (3,000 included) and throughput (125 MB/s included) set separately; aws ec2 modify-volume changes any of them online. The state goes modifying -> optimizing (the new size is usable from here) -> completed. Then, on the instance: growpart DISK N grows the partition, resize2fs (ext4) or xfs_growfs (XFS) grows the filesystem. One modification per volume every 6 hours, and a volume never shrinks. EC2 has no disk-usage metric: the CloudWatch agent publishes disk_used_percent.

What the instance sees

$ ssh -o StrictHostKeyChecking=accept-new ingest-1 'lsblk; df -h /'
Warning: Permanently added 'ingest-1' (ED25519) to the list of known hosts.
NAME         MAJ:MIN RM  SIZE RO TYPE MOUNTPOINTS
nvme0n1      259:0    0    8G  0 disk
├─nvme0n1p1  259:1    0    7G  0 part /
├─nvme0n1p14 259:2    0    4M  0 part
├─nvme0n1p15 259:3    0  106M  0 part /boot/efi
└─nvme0n1p16 259:4    0  913M  0 part /boot
Filesystem  Size  Used Avail Use% Mounted on
/dev/root   6.8G  4.6G  1.9G  72% /
$ aws ec2 describe-volumes --filters Name=attachment.instance-id,Values=i-0d4e5f60718293a41 --query 'Volumes[].{id: VolumeId, size: Size, type: VolumeType, iops: Iops, mbps: Throughput, device: Attachments[0].Device}'
[
    {
        "id": "vol-0f1e2d3c4b5a69788",
        "size": 8,
        "type": "gp3",
        "iops": 3000,
        "mbps": 125,
        "device": "/dev/sda1"
    }
]

Three layers, three sizes: the volume (8 GiB, what AWS bills), the partition nvme0n1p1 (7 GiB; Ubuntu's cloud image keeps the boot partitions p14, p15, p16 on the same disk), and the filesystem mounted on / (/dev/root, a bit smaller again: ext4 keeps room for its metadata). The volume is attached as /dev/sda1 in the EC2 API but appears as nvme0n1 inside the instance: on Nitro instances EBS volumes are NVMe devices, and the names do not match (ebsnvme-id or lsblk's SERIAL column, lsblk -o +SERIAL, maps them).

The disk metric that is not there

$ aws cloudwatch list-metrics --namespace AWS/EC2 --dimensions Name=InstanceId,Value=i-0d4e5f60718293a41 --query 'Metrics[].MetricName' --output text
CPUUtilization	EBSWriteBytes	NetworkIn	StatusCheckFailed
$ aws cloudwatch list-metrics --namespace CWAgent --query 'Metrics[].[MetricName,Dimensions[].[Name,Value][]]' --output json
[
    [
        "disk_used_percent",
        [
            "InstanceId",
            "i-0d4e5f60718293a41",
            "path",
            "/",
            "device",
            "nvme0n1p1",
            "fstype",
            "ext4"
        ]
    ],
    [
        "mem_used_percent",
        [
            "InstanceId",
            "i-0d4e5f60718293a41"
        ]
    ]
]

EC2 sees the instance from the hypervisor: CPU, network, EBS operations, status checks. It cannot see inside the filesystem - so disk space and memory come only from the CloudWatch agent (amazon-cloudwatch-agent), which publishes disk_used_percent with InstanceId, path, device and fstype dimensions - all four, exactly, in an alarm or a query:

$ aws cloudwatch get-metric-statistics --namespace CWAgent --metric-name disk_used_percent --dimensions Name=InstanceId,Value=i-0d4e5f60718293a41 Name=path,Value=/ Name=device,Value=nvme0n1p1 Name=fstype,Value=ext4 --start-time $(date -u -d '-6 hours' +%FT%TZ) --end-time $(date -u +%FT%TZ) --period 3600 --statistics Maximum --query 'sort_by(Datapoints, &Timestamp)[].[Timestamp,Maximum]' --output text
2026-09-22T14:00:00+00:00	60.1406
2026-09-22T15:00:00+00:00	62.3194
2026-09-22T16:00:00+00:00	64.4982
2026-09-22T17:00:00+00:00	66.677
2026-09-22T18:00:00+00:00	68.8557
2026-09-22T19:00:00+00:00	71
2026-09-22T20:00:00+00:00	71.0031

A steady climb, about two points an hour: at this rate / is full tomorrow morning. An alarm at 85% on this metric (with --treat-missing-data breaching: an agent that stopped reporting is also a problem) is the page you want - at 03:00 with hours to spare, not at 100%.

Growing the volume

$ aws ec2 modify-volume --volume-id vol-0f1e2d3c4b5a69788 --size 6
aws: [ERROR]: An error occurred (InvalidParameterValue) when calling the ModifyVolume operation: New size cannot be smaller than existing size
$ aws ec2 modify-volume --volume-id vol-0f1e2d3c4b5a69788 --size 12
{
    "VolumeModification": {
        "VolumeId": "vol-0f1e2d3c4b5a69788",
        "ModificationState": "modifying",
        "TargetSize": 12,
        "TargetIops": 3000,
        "TargetVolumeType": "gp3",
        "TargetThroughput": 125,
        "TargetMultiAttachEnabled": false,
        "OriginalSize": 8,
        "OriginalIops": 3000,
        "OriginalVolumeType": "gp3",
        "OriginalThroughput": 125,
        "OriginalMultiAttachEnabled": false,
        "Progress": 0,
        "StartTime": "2026-09-22T20:00:04.700+00:00"
    }
}
$ aws ec2 describe-volumes-modifications --volume-ids vol-0f1e2d3c4b5a69788 --query 'VolumesModifications[0].[ModificationState,OriginalSize,TargetSize,Progress]' --output text
modifying	8	12	0
$ sleep 10
$ aws ec2 describe-volumes-modifications --volume-ids vol-0f1e2d3c4b5a69788 --query 'VolumesModifications[0].[ModificationState,Progress]' --output text
optimizing	12

optimizing means the new size is already usable; the rest is EBS moving data in the background (performance can be a little uneven until completed). The instance sees a bigger disk - and still the same partition and filesystem:

$ ssh ingest-1 'lsblk /dev/nvme0n1; df -h /'
NAME         MAJ:MIN RM  SIZE RO TYPE MOUNTPOINTS
nvme0n1      259:0    0   12G  0 disk
├─nvme0n1p1  259:1    0    7G  0 part /
├─nvme0n1p14 259:2    0    4M  0 part
├─nvme0n1p15 259:3    0  106M  0 part /boot/efi
└─nvme0n1p16 259:4    0  913M  0 part /boot
Filesystem  Size  Used Avail Use% Mounted on
/dev/root   6.8G  4.6G  1.9G  72% /

Growing the partition and the filesystem

$ ssh ingest-1 growpart /dev/nvme0n1 1
failed [sfd_dump:1] sfdisk --unit=S --dump /dev/nvme0n1
sfdisk: cannot open /dev/nvme0n1: Permission denied
FAILED: failed to dump sfdisk info for /dev/nvme0n1
$ ssh ingest-1 sudo growpart /dev/nvme0n1 1
CHANGED: partition=1 start=2099200 old: size=14669824 end=16769023 new: size=23066591 end=25165790
$ ssh ingest-1 sudo resize2fs /dev/nvme0n1p1
resize2fs 1.47.2 (1-Jan-2025)
Filesystem at /dev/nvme0n1p1 is mounted on /; on-line resizing required
old_desc_blocks = 1, new_desc_blocks = 1
The filesystem on /dev/nvme0n1p1 is now 2883323 (4k) blocks long.
$ ssh ingest-1 'lsblk /dev/nvme0n1; df -h /'
NAME         MAJ:MIN RM  SIZE RO TYPE MOUNTPOINTS
nvme0n1      259:0    0   12G  0 disk
├─nvme0n1p1  259:1    0   11G  0 part /
├─nvme0n1p14 259:2    0    4M  0 part
├─nvme0n1p15 259:3    0  106M  0 part /boot/efi
└─nvme0n1p16 259:4    0  913M  0 part /boot
Filesystem  Size  Used Avail Use% Mounted on
/dev/root    11G  4.6G  5.6G  46% /

The two commands take two different arguments: growpart wants the disk and the partition number (/dev/nvme0n1 1 - not /dev/nvme0n1p1); resize2fs wants the partition (or /dev/root). Both run online on a mounted root filesystem, no reboot. On Amazon Linux the root filesystem is XFS: sudo xfs_growfs -d / instead of resize2fs. And on the next boot cloud-init's growpart module would have grown the root partition by itself - useful to know, but a reboot is exactly what you avoid at 03:00.

The 6-hour rule is the reason to grow generously the first time:

$ aws ec2 modify-volume --volume-id vol-0f1e2d3c4b5a69788 --size 16
aws: [ERROR]: An error occurred (VolumeModificationRateExceeded) when calling the ModifyVolume operation: You've reached the maximum modification rate per volume limit. Wait at least 6 hours between modifications per EBS volume.

gp3, gp2, io2 - and the snapshot first

Before a risky change to a disk - a filesystem resize on a database host, a type change - take a snapshot (aws ec2 create-snapshot --volume-id ..., incremental, stored in S3, a few cents per GB-month): it is the only undo. And a disk that fills up again needs the cause fixed, not another gigabyte: log rotation, retention, the job that writes too much.

In an interview: "The root volume of an EC2 instance is full and you cannot reboot it. What do you do?" - "Free a little space first if anything is safe to delete (old logs, journal vacuum), then aws ec2 modify-volume to a bigger size; once the modification is optimizing the instance sees the bigger disk, so sudo growpart /dev/nvme0n1 1 and sudo resize2fs on the partition (or xfs_growfs for XFS), online. Then fix why it filled up and add a CloudWatch agent disk alarm."

You can now: read a disk on an EC2 instance at its three layers (volume, partition, filesystem), explain why disk usage needs the CloudWatch agent, grow a gp3 volume online and watch the modification states, run growpart and resize2fs with the right arguments, and remember the 6-hour rule and the snapshot before risky disk changes.

Why it helps

"Disk full" is the oldest page in operations, and on an EC2 instance it has a calm, online fix if you know the three layers: the volume, the partition and the filesystem. Getting the order and the arguments right at 3 am is what this lesson rehearses.

It also explains why CloudWatch never warned you: EC2 cannot see disk space at all, so the warning has to come from the CloudWatch agent - and from an alarm that fires hours before the disk is full.

Commands in this lesson

ssh aws sleep

FAQ

Do I have to stop the instance to grow a volume?

No. aws ec2 modify-volume changes the size, type, IOPS or throughput of an attached volume online. Once the modification is in the optimizing state, the instance sees the bigger disk; then growpart and resize2fs (or xfs_growfs) grow the partition and the filesystem, also online.

Why does growpart fail with my partition name?

growpart takes the disk and the partition number as two arguments: growpart /dev/nvme0n1 1. Passing the partition itself (/dev/nvme0n1p1) makes it look for a partition table inside the partition and fail. resize2fs, on the other hand, takes the partition.

Can I shrink a volume or grow it twice in a night?

No to both. A volume can never be made smaller, and after a modification you must wait at least six hours before the next one on the same volume. Grow generously the first time, because the next fix may be hours away.

What is the difference between gp2 and gp3?

gp3 sets size, IOPS (3,000 included) and throughput (125 MB/s included) independently, so even a small volume is fast. gp2's performance grows with its size and small gp2 volumes run on burst credits that can run out under sustained load. Migrating to gp3 is an online modify-volume and usually cheaper.

Why is there no disk metric for my instance?

EC2 measures the instance from outside and cannot see the filesystem. The CloudWatch agent on the instance publishes disk_used_percent with InstanceId, path, device and fstype dimensions; alarms must use all four of them exactly, or they stay in INSUFFICIENT_DATA.

In an interview Mid

The root volume of an EC2 instance is full and you cannot reboot it. What do you do?

Free a little space if something is safe to delete, then aws ec2 modify-volume to a bigger size; once the modification is optimizing the instance sees the bigger disk, so sudo growpart /dev/nvme0n1 1 and sudo resize2fs on the partition (or xfs_growfs for XFS), online. Then fix why it filled up - rotation, retention - and add a CloudWatch agent disk alarm with lead time.

Also asked: Why does EC2 not publish disk usage, and how do you monitor it? · What changed between gp2 and gp3 volumes? · When would you take a snapshot before changing a volume?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.