OnCallReady

Lesson 31.32 · AWS III: CloudWatch, CloudTrail, Cost & Incidents · 15 min read

Well-Architected reliability: multi-AZ, backups and the incident on AWS

In plain words

Building something reliable on AWS is like planning a city that keeps running when a power station fails. You put things in more than one district, so losing one district does not stop the city. You keep copies of important records somewhere else, and you practise using them.

And when something does break, you follow the same steps every time: notice the alarm, look at the numbers, read the diaries, check what changed, fix it, and make sure it really is fixed.

Reliability on AWS: failure domains, backups, and the incident loop

The previous lessons were tools. This one is how they fit together when something breaks - and how a system is built so that fewer things break in the first place. AWS writes its advice down as the Well-Architected Framework; its reliability pillar is the part an SRE gets asked about, and its answers are concrete: assume every component fails, spread across Availability Zones, test your backups by restoring them, and know your RTO and RPO before the outage, not during it.

Need to know: failure domains nest - an instance, an AZ (one or more data centres), a Region, an account. Production runs in at least two AZs with health checks that take a bad one out. RPO = how much data you may lose (time since the last good copy), RTO = how long you may be down. A backup that was never restored is a hope. When it breaks, the loop is: alarm -> metrics (when, how much) -> logs (what) -> CloudTrail (what changed) -> mitigate -> verify - and first, "is it us or AWS?".

The Well-Architected Framework, briefly

Six pillars: operational excellence, security, reliability, performance efficiency, cost optimisation, sustainability. The reliability pillar's design principles, translated into what you would check in a review:

principlewhat it means in practice
recover automatically from failurehealth checks + Auto Scaling replace instances; alarms page a human only for what automation cannot fix
test recovery proceduresgame days; restore a backup into a scratch account every quarter
scale horizontallymany small instances behind a load balancer, not one big one
stop guessing capacityAuto Scaling on a metric; service quotas raised before launch day
manage change through automationIaC, pipelines, small deploys - the change CloudTrail shows should be a pipeline's

Failure domains and static stability

An AZ is one or more data centres with independent power and networking; AZ-wide failures are rare but real, and an AZ having a bad hour is the normal case to design for. So:

Bigger blast radius means a bigger domain: Regions for disaster recovery, accounts for "one mistake cannot touch everything" (Organizations, lesson 8).

Backups, RPO and RTO

strategyRPORTOcost
backup and restore (snapshots, AWS Backup, S3 copies)hourshours to a daylow
pilot light (data replicated, core infrastructure off)minutestens of minutesmedium
warm standby (a smaller full copy running)seconds to minutesminuteshigher
multi-site active/activenear zeronear zerohighest

The tools: EBS snapshots and AWS Backup plans (schedules, retention, copies to another Region or account, a vault lock against deletion), RDS automated backups and point-in-time restore, S3 versioning + replication (AWS I), DynamoDB point-in-time recovery. The two rules that matter more than the tool: keep a copy outside the blast radius (another Region, another account that the production admins cannot delete from), and restore regularly - a restore drill measures the real RTO.

Is it us or AWS?

When many things fail at once, check AWS's own status first - and know where that is:

$ aws health describe-events --region us-east-1 --filter eventStatusCodes=open
aws: [ERROR]: An error occurred (SubscriptionRequiredException) when calling the DescribeEvents operation: The AWS Premium Support Subscription is required to use this service.

The AWS Health API needs a Business, Enterprise On-Ramp or Enterprise support plan; on Basic or Developer support it answers SubscriptionRequiredException. Everyone has the AWS Health Dashboard in the console (account-specific events: an EC2 host retirement, an RDS maintenance, an issue in one AZ) and the public status page. Account-specific Health events can also go to EventBridge rules and on to SNS - free, and the way to get paged for "your instance is scheduled for retirement".

The incident loop, end to end

Checkout errors, half an hour ago. The tools of this chapter, in order:

$ aws cloudwatch describe-alarms --state-value ALARM --query 'MetricAlarms[].[AlarmName,StateTransitionedTimestamp]' --output text
try-shop-5xx	2026-09-22T19:36:07.000000+00:00
$ aws cloudwatch get-metric-statistics --namespace AWS/ApplicationELB --metric-name HTTPCode_Target_5XX_Count --dimensions Name=LoadBalancer,Value=app/shop-alb/50dc6c495c0c9188 --start-time $(date -u -d '-45 min' +%FT%TZ) --end-time $(date -u +%FT%TZ) --period 300 --statistics Sum --query 'sort_by(Datapoints, &Timestamp)[].[Timestamp,Sum]' --output text
2026-09-22T19:15:00+00:00	1
2026-09-22T19:20:00+00:00	1
2026-09-22T19:25:00+00:00	3
2026-09-22T19:30:00+00:00	97
2026-09-22T19:35:00+00:00	209
2026-09-22T19:40:00+00:00	213
2026-09-22T19:45:00+00:00	184
2026-09-22T19:50:00+00:00	235
2026-09-22T19:55:00+00:00	170
2026-09-22T20:00:00+00:00	36

When and how much: the alarm's transition time, and the metric's first bad period. Then what fails - the top errors since that time:

$ Q=$(aws logs start-query --log-group-name /oncall-lab/shop-api --start-time $(date -d '-45 min' +%s) --end-time $(date +%s) --query-string 'filter status >= 500 | stats count(*) as errors, min(@timestamp) as first by path, msg | sort errors desc' --query queryId --output text); sleep 2; aws logs get-query-results --query-id $Q --query 'results[*][*].value' --output text
/api/cart	no healthy upstream	55	2026-09-22 19:33:46.987
/api/cart/items	no healthy upstream	35	2026-09-22 19:32:10.050

And what changed just before - every write call in the window, in every Region that matters (us-east-1 for the global services):

$ for r in eu-central-1 us-east-1; do aws cloudtrail lookup-events --region $r --start-time $(date -u -d '-45 min' +%FT%TZ) --lookup-attributes AttributeKey=ReadOnly,AttributeValue=false --query 'Events[].[EventTime,EventSource,EventName,Username]' --output text; done
2026-09-22T20:00:04+00:00	logs.amazonaws.com	StartQuery	learner
2026-09-22T19:30:03+00:00	elasticloadbalancing.amazonaws.com	ModifyTargetGroup	gha-run-8812734
$ aws cloudtrail lookup-events --lookup-attributes AttributeKey=EventName,AttributeValue=ModifyTargetGroup --query 'Events[0].CloudTrailEvent' --output text | jq '{eventTime, who: .userIdentity.arn, agent: .userAgent, params: .requestParameters}'
{
  "eventTime": "2026-09-22T19:30:03Z",
  "who": "arn:aws:sts::111122223333:assumed-role/gha-shop-infra/gha-run-8812734",
  "agent": "APN/1.0 HashiCorp/1.0 Terraform/1.16.5 (+https://www.terraform.io) terraform-provider-aws/6.21.0 (+https://registry.terraform.io/providers/hashicorp/aws) aws-sdk-go-v2/1.39.6 os/linux lang/go#1.25.3 md/GOOS#linux md/GOARCH#amd64",
  "params": {
    "targetGroupArn": "arn:aws:elasticloadbalancing:eu-central-1:111122223333:targetgroup/shop-web/6d0ecf831eec9f09",
    "healthCheckPath": "/healthz",
    "healthCheckIntervalSeconds": 10
  }
}

Two minutes before the first error, a pipeline run changed the target group's health check path. That is the hypothesis: a health check against a path that does not exist marks targets unhealthy, the load balancer has fewer (or no) healthy targets, requests fail. Mitigate first (roll the change back - here, the health check path), verify with the same metric that showed the problem, then find out why the pipeline shipped it. Writing the timeline as you go - alarm time, first error, the change, the mitigation, recovery - is what makes the post-incident review possible.

Game days and chaos

The only way to know the system survives an AZ failure is to cause one on purpose, in a controlled way: a game day with a hypothesis ("if 1a's instances stop, checkout keeps working and the alarm pages within 3 minutes"), a stop button, and notes. AWS Fault Injection Service (FIS) runs such experiments - stop instances, inject API errors, throttle, simulate an AZ power interruption - with stop conditions tied to CloudWatch alarms. Start in staging; graduate to production once the alarms and runbooks have proven themselves.

In an interview: "How would you make a web service on AWS highly available?" - "Run it in at least two Availability Zones behind a load balancer with health checks, in an Auto Scaling group sized so the surviving AZs carry the load if one fails, with managed data stores in Multi-AZ mode and backups copied to another Region or account and restored regularly. Then alarms on user-facing symptoms, and game days to prove it - including RTO and RPO targets agreed upfront."

You can now: name the reliability pillar's principles and turn them into review questions, reason about failure domains and static stability, pick a DR strategy from RTO and RPO, say where AWS's own status is (and what the Health API needs), and run the incident loop - alarm, metrics, logs, CloudTrail, mitigate, verify - with the commands from this chapter.

Why it helps

Interviews and design reviews ask "how would you make this highly available?", and the answer has a standard shape: multiple Availability Zones, static stability, backups you have restored, RTO and RPO agreed upfront, and alarms on what users feel.

The incident loop at the end of the lesson ties the whole chapter together: alarm, metrics, logs, CloudTrail, mitigate, verify. It is the habit that makes the difference between a calm incident and a long one.

Commands in this lesson

aws

FAQ

What is an Availability Zone?

One or more data centres in a Region with independent power, cooling and networking, connected to the other zones by fast private links. Zone-wide failures are rare but real; designing for the loss of one zone is the standard for production workloads.

What do RTO and RPO mean?

RPO, the recovery point objective, is how much data you can afford to lose - the time since the last good copy. RTO, the recovery time objective, is how long you can afford to be down. Together they choose the disaster-recovery strategy and its cost.

How do I check whether AWS itself has a problem?

The AWS Health Dashboard in the console shows events that affect your account, and the public status page shows the rest. The AWS Health API needs a Business, Enterprise On-Ramp or Enterprise support plan; with Basic support it answers SubscriptionRequiredException.

What is static stability?

Running with enough capacity in the remaining zones that losing one zone needs no new launches. During a wide outage the control plane that launches replacements is busy too, so a system that must scale up to survive a failure may not get the capacity in time.

What is a game day?

A planned exercise where you cause a failure on purpose - stop instances, cut a dependency, simulate a zone outage - with a hypothesis, a stop button and notes, to prove alarms, runbooks and capacity work. AWS Fault Injection Service runs such experiments with stop conditions tied to CloudWatch alarms.

In an interview Mid

How would you make a web service on AWS highly available?

Run it in at least two Availability Zones behind a load balancer with health checks, in an Auto Scaling group sized so the surviving zones carry the load if one fails (static stability), with managed data stores in Multi-AZ mode and backups copied to another Region or account and restored regularly. Then alarms on user-facing symptoms, agreed RTO and RPO targets, and game days to prove it.

Also asked: What is the difference between RTO and RPO? · How do you tell whether an outage is yours or AWS's? · Walk me through how you would work an incident on AWS.

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.