OnCallReady

Lesson 30.24 · AWS II: VPC, EC2, ELB & EKS · 12 min read

Auto Scaling groups: desired capacity, health, instance refresh

In plain words

An Auto Scaling group is a shift manager with a staffing number on a whiteboard. If a worker goes home sick, the manager calls in a replacement from the same job description (the launch template). If the whiteboard says more, the manager hires more; if less, sends some home. To change everyone's uniform, the manager swaps workers a few at a time so the shop never closes - that is an instance refresh.

Auto Scaling groups: desired capacity, health, instance refresh

Hand-launched instances are pets: someone remembers how they were made. An Auto Scaling group (ASG) turns them into cattle: a recipe (the launch template), a number (desired capacity), and a loop that keeps the real world equal to the number - replacing what dies, spreading over AZs, and rolling new versions out. EKS managed node groups are ASGs too, so this lesson is also how your Kubernetes nodes behave.

Need to know: an ASG keeps desired instances (between min and max) from a launch template version, spread over its subnets' AZs. With EC2 health checks it replaces instances whose status checks fail; with ELB health checks it also replaces the ones its target group calls unhealthy, after the grace period. Every change is a scaling activity with a cause. A new image ships as a new launch template version plus an instance refresh, which replaces instances in batches while keeping MinHealthyPercentage of the capacity in service.

The group

$ aws autoscaling describe-auto-scaling-groups --auto-scaling-group-names try-asg --query 'AutoScalingGroups[0].[MinSize,MaxSize,DesiredCapacity,HealthCheckType,HealthCheckGracePeriod,LaunchTemplate.LaunchTemplateName,LaunchTemplate.Version]' --output text
2	4	2	EC2	300	try-web	$Latest
$ aws autoscaling describe-auto-scaling-groups --auto-scaling-group-names try-asg --query 'AutoScalingGroups[0].Instances[].[InstanceId,AvailabilityZone,LifecycleState,HealthStatus,LaunchTemplate.Version]' --output table
-----------------------------------------------------------------------
|                      DescribeAutoScalingGroups                      |
+----------------------+----------------+------------+----------+-----+
|  i-0859d7e46b1dd17e3 |  eu-central-1a |  InService |  Healthy |  1  |
|  i-0af9190646233a59a |  eu-central-1b |  InService |  Healthy |  1  |
+----------------------+----------------+------------+----------+-----+

Two instances, one per AZ (the group balances across its subnets' AZs), both InService. The template version is $Latest: every new instance uses whatever the newest version is when it launches. $Default (the version marked default) or a number pins it.

The diary: scaling activities

$ aws autoscaling describe-scaling-activities --auto-scaling-group-name try-asg --max-items 3 --query 'Activities[].[StatusCode,Description,Cause]' --output text
Successful	Launching a new EC2 instance: i-0af9190646233a59a	At 2026-09-22T19:40:03Z an instance was started in response to a difference between desired and actual capacity, increasing the capacity from 1 to 2.
Successful	Launching a new EC2 instance: i-0859d7e46b1dd17e3	At 2026-09-22T19:40:03Z an instance was started in response to a difference between desired and actual capacity, increasing the capacity from 0 to 1.

Every launch and termination has a description and a cause - "a user request update of AutoScalingGroup constraints to min: 2, max: 4, desired: 2 changing the desired capacity from 0 to 2", "an instance was taken out of service in response to an ELB system health check failure". When instances appear or vanish and nobody knows why, this is the first command.

The loop at work

Terminate an instance behind the group's back and watch it come back:

$ VICTIM=$(aws autoscaling describe-auto-scaling-groups --auto-scaling-group-names try-asg --query 'AutoScalingGroups[0].Instances[0].InstanceId' --output text)
$ aws ec2 terminate-instances --instance-ids $VICTIM --query 'TerminatingInstances[0].CurrentState.Name' --output text
shutting-down
$ sleep 60
$ aws autoscaling describe-scaling-activities --auto-scaling-group-name try-asg --max-items 2 --query 'Activities[].[StatusCode,Description]' --output text
Successful	Launching a new EC2 instance: i-0f40119695729a52a
Successful	Terminating EC2 instance: i-0859d7e46b1dd17e3

That is the point of an ASG - and the classic surprise: you cannot fix a broken instance by terminating it if the template is what is broken; the group launches the same broken thing again. To take an instance out for good, terminate-instance-in-auto-scaling-group --should-decrement-desired-capacity, or change desired.

Health: EC2 or ELB

Health check typeReplaces an instance when
EC2 (default)its EC2 status checks fail (the VM or the host is broken)
ELBalso when a target group it is attached to calls it unhealthy

With ELB the grace period (seconds after launch) keeps the group from killing instances that are still booting; set it to how long your app really takes to become healthy. Too short and a slow-starting app is replaced forever (a "launch, fail, terminate" loop in the activities); too long and dead instances serve errors for minutes.

Scaling policies (briefly)

Desired can be changed by hand (set-desired-capacity), on a schedule, or by a target tracking policy - "keep average CPU at 50%" or "keep requests per target at 1,000" - which raises and lowers desired for you, within min and max. Max is the safety net for your bill.

Releasing a new version: instance refresh

$ aws ec2 create-launch-template-version --launch-template-name try-web --source-version 1 --version-description 'shop-web 2.4.1' --launch-template-data '{"ImageId":"'$(aws ec2 describe-images --owners self --filters 'Name=name,Values=oncall-web-*' --query 'sort_by(Images, &CreationDate)[-1].ImageId' --output text)'"}' --query 'LaunchTemplateVersion.[VersionNumber,VersionDescription]' --output text
2	shop-web 2.4.1
$ aws autoscaling start-instance-refresh --auto-scaling-group-name try-asg --preferences '{"MinHealthyPercentage":50,"InstanceWarmup":60}'
{
    "InstanceRefreshId": "25175a0d-3824-4d21-9b24-730b882868ad"
}
$ sleep 30
$ aws autoscaling describe-instance-refreshes --auto-scaling-group-name try-asg --query 'InstanceRefreshes[0].[Status,PercentageComplete,InstancesToUpdate,StatusReason]' --output text
InProgress	0	2	Waiting for instances to warm up before continuing. For example: i-0c00283619d0e556e is warming up.

An instance refresh terminates and replaces instances in batches: with 2 instances and MinHealthyPercentage 50, one at a time. Each new instance must pass its health checks and then InstanceWarmup seconds before the next batch. If new instances never become healthy, the refresh stops at its checkpoint and fails - with AutoRollback it goes back to the previous template version. Two more preferences worth knowing: SkipMatching (do not replace instances already on the desired version) and checkpoints (pause at 50% for a manual check).

In an interview: "How do you roll out a new AMI to an Auto Scaling group without downtime?" - a new launch template version with the new image, then an instance refresh with a MinHealthyPercentage and a warmup so batches are replaced only after the new instances pass the load balancer's health checks (and auto-rollback if they never do).

What you can do now

Why it helps

Groups decide what runs, so they explain the strangest on-call moments: an instance that comes back after you terminated it, a fleet stuck replacing instances that never become healthy, a release that took half the capacity down. Reading activities, health check settings and refresh status lets you see what the group is doing and why, and ship new images safely.

Commands in this lesson

aws sleep

FAQ

Why did my terminated instance come back?

Terminating an instance does not change the group's desired capacity, so the group sees one instance missing and launches a replacement from the same launch template. To remove an instance for good, terminate it through Auto Scaling with should-decrement-desired-capacity, or lower desired. If the template is broken, fix the template.

EC2 or ELB health checks?

EC2 health checks only replace instances whose status checks fail, so an instance whose application is dead stays in service. With ELB health checks the group also replaces instances the target group calls unhealthy, after the grace period. Behind a load balancer, use ELB and set the grace period to how long the app really needs to start.

What is the grace period for?

It gives new instances time to boot and start the application before ELB health checks count against them. Too short and slow instances are replaced in a loop, visible as launch and terminate pairs in the scaling activities. Too long and a dead instance serves errors for minutes before it is replaced.

How does an instance refresh avoid downtime?

It replaces instances in batches and never goes below MinHealthyPercentage of the capacity in service. Each new instance must pass its health checks and then wait the warmup before the next batch. If new instances never become healthy, the refresh stops; with AutoRollback it returns to the previous launch template version.

Where do I see what the group did?

In describe-scaling-activities: every launch and termination with a status, a description and a cause, such as a change in desired capacity, an unhealthy instance or an instance refresh. It is the first command when instances appear or disappear and nobody knows why.

In an interview Mid

How do you roll out a new AMI to an Auto Scaling group without downtime?

Create a new launch template version with the new image (create-launch-template-version --source-version keeps the rest), then start-instance-refresh with a MinHealthyPercentage and an InstanceWarmup: instances are replaced in batches, each new one has to pass the load balancer's health checks and warm up before the next batch, and capacity never drops below the minimum. With AutoRollback a refresh whose new instances never become healthy goes back to the previous version.

Also asked: What is the difference between min, max and desired capacity? · Why would an Auto Scaling group keep replacing healthy-looking instances? · How does target tracking scaling work?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.