OnCallReady

Chapter 31 AWS III: CloudWatch, CloudTrail, Cost & Incidents

Running a workload on AWS when it breaks: CloudWatch metrics and alarms that page you through SNS, CloudWatch Logs with filter patterns and Logs Insights queries, CloudTrail to find out who changed what, Cost Explorer and the bills that surprise teams (the NAT gateway), tags, budgets and cost allocation, Organizations and service control policies, growing a full EBS volume online, and an outage worked end to end - against a simulated account with real outputs and real errors.

In plain words

Imagine running a shop in a rented building. You need a few things to sleep well: gauges on the wall that show how busy the tills are and how many sales fail, a bell that rings when a gauge goes into the red, a notebook of everything that happened, a camera that records who opened which door, and a monthly bill you can actually read.

On AWS those are CloudWatch (the gauges and the bell), CloudWatch Logs (the notebook), CloudTrail (the camera) and Cost Explorer (the bill). This chapter teaches you to read all four when something goes wrong.

Why it matters on call

Building on AWS is the easy part; being on call for it is the job. Most AWS incidents are answered with the same four tools: a metric that shows when it started, a log query that shows what fails, a CloudTrail record that shows what changed and who changed it, and a cost breakdown when the surprise is on the invoice instead of the dashboard.

The chapter also covers the guardrails around an account - tags and budgets, Organizations and service control policies - and the disk that fills up on an instance you cannot reboot, so the first time you meet these is in a lab, not at 3 am.

Lessons

  1. CloudWatch metrics: namespaces, dimensions, statistics and periods
  2. Alarms: states, M out of N, missing data and notifications
  3. CloudWatch Logs: groups, retention, filter patterns and tail
  4. Logs Insights: query your logs like a database
  5. CloudTrail: event history, trails and reading a record
  6. The cost model: what you pay for, and the bills that surprise people
  7. Tags, cost allocation and budgets
  8. Organizations and service control policies
  9. EBS from the instance: gp3, growing a disk online, and the disk metrics EC2 does not have
  10. Well-Architected reliability: multi-AZ, backups and the incident on AWS

23 hands-on labs (missions, incidents and drills) run in the terminal: Open this chapter in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.

Questions people ask

Do I need AWS II before this chapter?

It helps but it is not required. The chapter runs in the same simulated account as AWS I and II and brings its own small shop: a load balancer, a few instances, a NAT gateway, log groups and a bill. Where a lab needs something AWS II teaches (a VPC endpoint, an EBS volume), the lesson explains the parts you use.

Is CloudWatch the same as Prometheus or Azure Monitor?

It plays the same role - metrics, alarms and logs for everything in the account - but every service publishes into it by itself, with no agent to install for the basics. The ideas transfer: a metric has a name and labels (called dimensions), an alarm watches a threshold over time, and logs are queried with a small query language that looks a lot like Azure's KQL.

Will the labs cost money if I repeat them on a real account?

Most of what the labs do costs cents in a real account: a few alarms, a log group, some Cost Explorer queries (each Cost Explorer API call is billed at one cent). The expensive part of the chapter is what it warns about - NAT gateway traffic, forgotten instances, logs without retention - so set a budget before you start experimenting.

Why does the chapter spend so much time on cost?

Because a cost spike is an incident that nobody pages for. The bill is a monitoring signal like latency, and the investigation is the same: find when it changed, what changed, and who owns it. Teams that can explain their bill line by line get trusted with bigger budgets; the NAT gateway incident in this chapter is one of the most common real ones.

What is simulated and what is real in this chapter?

The commands, their options, the output formats and the error messages follow the real AWS CLI v2 and the current documentation. The account behind them is simulated (simulator): metrics are computed from deterministic patterns, logs are generated, and the shop's instances are lab hosts. Everything you type works the same way against a real account.