Auto Scaling groups: desired capacity, health, instance refresh
Hand-launched instances are pets: someone remembers how they were made. An Auto Scaling group (ASG) turns them into cattle: a recipe (the launch template), a number (desired capacity), and a loop that keeps the real world equal to the number - replacing what dies, spreading over AZs, and rolling new versions out. EKS managed node groups are ASGs too, so this lesson is also how your Kubernetes nodes behave.
Need to know: an ASG keeps desired instances (between min and max) from a launch template version, spread over its subnets' AZs. With EC2 health checks it replaces instances whose status checks fail; with ELB health checks it also replaces the ones its target group calls unhealthy, after the grace period. Every change is a scaling activity with a cause. A new image ships as a new launch template version plus an instance refresh, which replaces instances in batches while keeping MinHealthyPercentage of the capacity in service.
The group
$ aws autoscaling describe-auto-scaling-groups --auto-scaling-group-names try-asg --query 'AutoScalingGroups[0].[MinSize,MaxSize,DesiredCapacity,HealthCheckType,HealthCheckGracePeriod,LaunchTemplate.LaunchTemplateName,LaunchTemplate.Version]' --output text
2 4 2 EC2 300 try-web $Latest
$ aws autoscaling describe-auto-scaling-groups --auto-scaling-group-names try-asg --query 'AutoScalingGroups[0].Instances[].[InstanceId,AvailabilityZone,LifecycleState,HealthStatus,LaunchTemplate.Version]' --output table
-----------------------------------------------------------------------
| DescribeAutoScalingGroups |
+----------------------+----------------+------------+----------+-----+
| i-0859d7e46b1dd17e3 | eu-central-1a | InService | Healthy | 1 |
| i-0af9190646233a59a | eu-central-1b | InService | Healthy | 1 |
+----------------------+----------------+------------+----------+-----+
Two instances, one per AZ (the group balances across its subnets' AZs), both InService. The template version is $Latest: every new instance uses whatever the newest version is when it launches. $Default (the version marked default) or a number pins it.
The diary: scaling activities
$ aws autoscaling describe-scaling-activities --auto-scaling-group-name try-asg --max-items 3 --query 'Activities[].[StatusCode,Description,Cause]' --output text
Successful Launching a new EC2 instance: i-0af9190646233a59a At 2026-09-22T19:40:03Z an instance was started in response to a difference between desired and actual capacity, increasing the capacity from 1 to 2.
Successful Launching a new EC2 instance: i-0859d7e46b1dd17e3 At 2026-09-22T19:40:03Z an instance was started in response to a difference between desired and actual capacity, increasing the capacity from 0 to 1.
Every launch and termination has a description and a cause - "a user request update of AutoScalingGroup constraints to min: 2, max: 4, desired: 2 changing the desired capacity from 0 to 2", "an instance was taken out of service in response to an ELB system health check failure". When instances appear or vanish and nobody knows why, this is the first command.
The loop at work
Terminate an instance behind the group's back and watch it come back:
$ VICTIM=$(aws autoscaling describe-auto-scaling-groups --auto-scaling-group-names try-asg --query 'AutoScalingGroups[0].Instances[0].InstanceId' --output text)
$ aws ec2 terminate-instances --instance-ids $VICTIM --query 'TerminatingInstances[0].CurrentState.Name' --output text
shutting-down
$ sleep 60
$ aws autoscaling describe-scaling-activities --auto-scaling-group-name try-asg --max-items 2 --query 'Activities[].[StatusCode,Description]' --output text
Successful Launching a new EC2 instance: i-0f40119695729a52a
Successful Terminating EC2 instance: i-0859d7e46b1dd17e3
That is the point of an ASG - and the classic surprise: you cannot fix a broken instance by terminating it if the template is what is broken; the group launches the same broken thing again. To take an instance out for good, terminate-instance-in-auto-scaling-group --should-decrement-desired-capacity, or change desired.
Health: EC2 or ELB
| Health check type | Replaces an instance when |
|---|---|
EC2 (default) | its EC2 status checks fail (the VM or the host is broken) |
ELB | also when a target group it is attached to calls it unhealthy |
With ELB the grace period (seconds after launch) keeps the group from killing instances that are still booting; set it to how long your app really takes to become healthy. Too short and a slow-starting app is replaced forever (a "launch, fail, terminate" loop in the activities); too long and dead instances serve errors for minutes.
Scaling policies (briefly)
Desired can be changed by hand (set-desired-capacity), on a schedule, or by a target tracking policy - "keep average CPU at 50%" or "keep requests per target at 1,000" - which raises and lowers desired for you, within min and max. Max is the safety net for your bill.
Releasing a new version: instance refresh
$ aws ec2 create-launch-template-version --launch-template-name try-web --source-version 1 --version-description 'shop-web 2.4.1' --launch-template-data '{"ImageId":"'$(aws ec2 describe-images --owners self --filters 'Name=name,Values=oncall-web-*' --query 'sort_by(Images, &CreationDate)[-1].ImageId' --output text)'"}' --query 'LaunchTemplateVersion.[VersionNumber,VersionDescription]' --output text
2 shop-web 2.4.1
$ aws autoscaling start-instance-refresh --auto-scaling-group-name try-asg --preferences '{"MinHealthyPercentage":50,"InstanceWarmup":60}'
{
"InstanceRefreshId": "25175a0d-3824-4d21-9b24-730b882868ad"
}
$ sleep 30
$ aws autoscaling describe-instance-refreshes --auto-scaling-group-name try-asg --query 'InstanceRefreshes[0].[Status,PercentageComplete,InstancesToUpdate,StatusReason]' --output text
InProgress 0 2 Waiting for instances to warm up before continuing. For example: i-0c00283619d0e556e is warming up.
An instance refresh terminates and replaces instances in batches: with 2 instances and MinHealthyPercentage 50, one at a time. Each new instance must pass its health checks and then InstanceWarmup seconds before the next batch. If new instances never become healthy, the refresh stops at its checkpoint and fails - with AutoRollback it goes back to the previous template version. Two more preferences worth knowing: SkipMatching (do not replace instances already on the desired version) and checkpoints (pause at 50% for a manual check).
In an interview: "How do you roll out a new AMI to an Auto Scaling group without downtime?" - a new launch template version with the new image, then an instance refresh with a MinHealthyPercentage and a warmup so batches are replaced only after the new instances pass the load balancer's health checks (and auto-rollback if they never do).
What you can do now
- read a group's size, health settings and instances, and its activity history;
- predict what the group does when an instance dies, is unhealthy, or is terminated by hand;
- ship a new image with a launch template version and an instance refresh.