Reliability on AWS: failure domains, backups, and the incident loop
The previous lessons were tools. This one is how they fit together when something breaks - and how a system is built so that fewer things break in the first place. AWS writes its advice down as the Well-Architected Framework; its reliability pillar is the part an SRE gets asked about, and its answers are concrete: assume every component fails, spread across Availability Zones, test your backups by restoring them, and know your RTO and RPO before the outage, not during it.
Need to know: failure domains nest - an instance, an AZ (one or more data centres), a Region, an account. Production runs in at least two AZs with health checks that take a bad one out. RPO = how much data you may lose (time since the last good copy), RTO = how long you may be down. A backup that was never restored is a hope. When it breaks, the loop is: alarm -> metrics (when, how much) -> logs (what) -> CloudTrail (what changed) -> mitigate -> verify - and first, "is it us or AWS?".
The Well-Architected Framework, briefly
Six pillars: operational excellence, security, reliability, performance efficiency, cost optimisation, sustainability. The reliability pillar's design principles, translated into what you would check in a review:
| principle | what it means in practice |
|---|---|
| recover automatically from failure | health checks + Auto Scaling replace instances; alarms page a human only for what automation cannot fix |
| test recovery procedures | game days; restore a backup into a scratch account every quarter |
| scale horizontally | many small instances behind a load balancer, not one big one |
| stop guessing capacity | Auto Scaling on a metric; service quotas raised before launch day |
| manage change through automation | IaC, pipelines, small deploys - the change CloudTrail shows should be a pipeline's |
Failure domains and static stability
An AZ is one or more data centres with independent power and networking; AZ-wide failures are rare but real, and an AZ having a bad hour is the normal case to design for. So:
- spread instances over two or three AZs, behind a load balancer whose health checks stop sending traffic to bad targets (and turn on cross-zone load balancing deliberately);
- managed services offer it as a switch: RDS Multi-AZ (a synchronous standby, failover in about a minute), EFS, S3 (multi-AZ by default), DynamoDB;
- design for static stability: if an AZ fails, the remaining AZs already have enough capacity (N+1), because launching replacements during a widespread event may be slow - the control plane is busy too;
- one NAT gateway per AZ, not one for the VPC: the shop's single NAT gateway in 1a is both a cost problem (cross-AZ bytes, the incident in this chapter) and an availability one (1a down = 1b's private subnets lose the internet).
Bigger blast radius means a bigger domain: Regions for disaster recovery, accounts for "one mistake cannot touch everything" (Organizations, lesson 8).
Backups, RPO and RTO
| strategy | RPO | RTO | cost |
|---|---|---|---|
| backup and restore (snapshots, AWS Backup, S3 copies) | hours | hours to a day | low |
| pilot light (data replicated, core infrastructure off) | minutes | tens of minutes | medium |
| warm standby (a smaller full copy running) | seconds to minutes | minutes | higher |
| multi-site active/active | near zero | near zero | highest |
The tools: EBS snapshots and AWS Backup plans (schedules, retention, copies to another Region or account, a vault lock against deletion), RDS automated backups and point-in-time restore, S3 versioning + replication (AWS I), DynamoDB point-in-time recovery. The two rules that matter more than the tool: keep a copy outside the blast radius (another Region, another account that the production admins cannot delete from), and restore regularly - a restore drill measures the real RTO.
Is it us or AWS?
When many things fail at once, check AWS's own status first - and know where that is:
$ aws health describe-events --region us-east-1 --filter eventStatusCodes=open
aws: [ERROR]: An error occurred (SubscriptionRequiredException) when calling the DescribeEvents operation: The AWS Premium Support Subscription is required to use this service.
The AWS Health API needs a Business, Enterprise On-Ramp or Enterprise support plan; on Basic or Developer support it answers SubscriptionRequiredException. Everyone has the AWS Health Dashboard in the console (account-specific events: an EC2 host retirement, an RDS maintenance, an issue in one AZ) and the public status page. Account-specific Health events can also go to EventBridge rules and on to SNS - free, and the way to get paged for "your instance is scheduled for retirement".
The incident loop, end to end
Checkout errors, half an hour ago. The tools of this chapter, in order:
$ aws cloudwatch describe-alarms --state-value ALARM --query 'MetricAlarms[].[AlarmName,StateTransitionedTimestamp]' --output text
try-shop-5xx 2026-09-22T19:36:07.000000+00:00
$ aws cloudwatch get-metric-statistics --namespace AWS/ApplicationELB --metric-name HTTPCode_Target_5XX_Count --dimensions Name=LoadBalancer,Value=app/shop-alb/50dc6c495c0c9188 --start-time $(date -u -d '-45 min' +%FT%TZ) --end-time $(date -u +%FT%TZ) --period 300 --statistics Sum --query 'sort_by(Datapoints, &Timestamp)[].[Timestamp,Sum]' --output text
2026-09-22T19:15:00+00:00 1
2026-09-22T19:20:00+00:00 1
2026-09-22T19:25:00+00:00 3
2026-09-22T19:30:00+00:00 97
2026-09-22T19:35:00+00:00 209
2026-09-22T19:40:00+00:00 213
2026-09-22T19:45:00+00:00 184
2026-09-22T19:50:00+00:00 235
2026-09-22T19:55:00+00:00 170
2026-09-22T20:00:00+00:00 36
When and how much: the alarm's transition time, and the metric's first bad period. Then what fails - the top errors since that time:
$ Q=$(aws logs start-query --log-group-name /oncall-lab/shop-api --start-time $(date -d '-45 min' +%s) --end-time $(date +%s) --query-string 'filter status >= 500 | stats count(*) as errors, min(@timestamp) as first by path, msg | sort errors desc' --query queryId --output text); sleep 2; aws logs get-query-results --query-id $Q --query 'results[*][*].value' --output text
/api/cart no healthy upstream 55 2026-09-22 19:33:46.987
/api/cart/items no healthy upstream 35 2026-09-22 19:32:10.050
And what changed just before - every write call in the window, in every Region that matters (us-east-1 for the global services):
$ for r in eu-central-1 us-east-1; do aws cloudtrail lookup-events --region $r --start-time $(date -u -d '-45 min' +%FT%TZ) --lookup-attributes AttributeKey=ReadOnly,AttributeValue=false --query 'Events[].[EventTime,EventSource,EventName,Username]' --output text; done
2026-09-22T20:00:04+00:00 logs.amazonaws.com StartQuery learner
2026-09-22T19:30:03+00:00 elasticloadbalancing.amazonaws.com ModifyTargetGroup gha-run-8812734
$ aws cloudtrail lookup-events --lookup-attributes AttributeKey=EventName,AttributeValue=ModifyTargetGroup --query 'Events[0].CloudTrailEvent' --output text | jq '{eventTime, who: .userIdentity.arn, agent: .userAgent, params: .requestParameters}'
{
"eventTime": "2026-09-22T19:30:03Z",
"who": "arn:aws:sts::111122223333:assumed-role/gha-shop-infra/gha-run-8812734",
"agent": "APN/1.0 HashiCorp/1.0 Terraform/1.16.5 (+https://www.terraform.io) terraform-provider-aws/6.21.0 (+https://registry.terraform.io/providers/hashicorp/aws) aws-sdk-go-v2/1.39.6 os/linux lang/go#1.25.3 md/GOOS#linux md/GOARCH#amd64",
"params": {
"targetGroupArn": "arn:aws:elasticloadbalancing:eu-central-1:111122223333:targetgroup/shop-web/6d0ecf831eec9f09",
"healthCheckPath": "/healthz",
"healthCheckIntervalSeconds": 10
}
}
Two minutes before the first error, a pipeline run changed the target group's health check path. That is the hypothesis: a health check against a path that does not exist marks targets unhealthy, the load balancer has fewer (or no) healthy targets, requests fail. Mitigate first (roll the change back - here, the health check path), verify with the same metric that showed the problem, then find out why the pipeline shipped it. Writing the timeline as you go - alarm time, first error, the change, the mitigation, recovery - is what makes the post-incident review possible.
Game days and chaos
The only way to know the system survives an AZ failure is to cause one on purpose, in a controlled way: a game day with a hypothesis ("if 1a's instances stop, checkout keeps working and the alarm pages within 3 minutes"), a stop button, and notes. AWS Fault Injection Service (FIS) runs such experiments - stop instances, inject API errors, throttle, simulate an AZ power interruption - with stop conditions tied to CloudWatch alarms. Start in staging; graduate to production once the alarms and runbooks have proven themselves.
In an interview: "How would you make a web service on AWS highly available?" - "Run it in at least two Availability Zones behind a load balancer with health checks, in an Auto Scaling group sized so the surviving AZs carry the load if one fails, with managed data stores in Multi-AZ mode and backups copied to another Region or account and restored regularly. Then alarms on user-facing symptoms, and game days to prove it - including RTO and RPO targets agreed upfront."
You can now: name the reliability pillar's principles and turn them into review questions, reason about failure domains and static stability, pick a DR strategy from RTO and RPO, say where AWS's own status is (and what the Health API needs), and run the incident loop - alarm, metrics, logs, CloudTrail, mitigate, verify - with the commands from this chapter.