AWS III: CloudWatch, CloudTrail, Cost & Incidents: interview questions
The question you are most likely to get for each topic, a model answer, and what else comes up. From chapter 31 of the course.
Checkout errors start at 3 am on AWS. Walk me through your first fifteen minutes. Mid
Start from the page: the alarm's history and the metric tell when it started and how bad it is. Then Logs Insights on the service's log group - filter status >= 500 | stats count(*) by msg and a bin(5m) timeline - says what fails. Then CloudTrail lookup-events for write calls just before the start, in the workload's Region and us-east-1, says what changed. Mitigate (roll back), verify with the same metric, and keep a timeline as you go.
Also asked: How do you find out who deleted a resource in an AWS account? · The AWS bill doubled this month. How do you explain it? · What is the difference between an IAM policy and a service control policy?
Average latency looks fine but users complain about slowness. What do you look at? Mid
The percentiles: p99 or p95 of the load balancer's TargetResponseTime with get-metric-statistics --extended-statistics p99. An average of 60 ms can hide a p99 of several seconds - one request in a hundred is slow, and those are the users who complain. Then break it down by target or path to find who is slow, and alarm on the percentile rather than the average.
Also asked: Why might a CloudWatch query return no data for a metric you know exists? · How do you compute an error rate from two CloudWatch metrics? · What metrics does EC2 not publish, and how do you get them?
Learn it: 31.1 CloudWatch metrics: namespaces, dimensions, statistics and periods
An alarm never fired during an outage. How do you find out why? Mid
Read its history and StateReasonData with describe-alarm-history: was it sitting in INSUFFICIENT_DATA because missing data was left as "missing" on a sparse metric, were the dimensions exactly the metric's, was M out of N too strict for the period? Then check the action path: the SNS topic exists and the subscription is confirmed, not PendingConfirmation. Finally test the whole path with set-alarm-state.
Also asked: When would you treat missing data as breaching? · Why alarm on an error rate instead of an error count? · How do you stop an alarm from paging during a maintenance window?
Learn it: 31.4 Alarms: states, M out of N, missing data and notifications
How would you get alerted on a specific error message in an application's logs on AWS? Mid
Create a metric filter on the application's log group with put-metric-filter: a filter pattern for the message (a JSON selector if the logs are structured), publishing a count with defaultValue 0. Then a CloudWatch alarm on that metric with an SNS action. A metric filter only counts new events, so test the pattern with test-metric-filter first.
Also asked: Why is "Never expire" a problem for CloudWatch Logs? · How do you follow the logs of a whole fleet of instances live? · What is the difference between a metric filter and a subscription filter?
Learn it: 31.8 CloudWatch Logs: groups, retention, filter patterns and tail
Users report errors since this morning. How do you find out what is failing from the logs? Mid
A Logs Insights query over the service's log group: filter status >= 500 (or level = ERROR) | stats count(*) by msg, path | sort desc to see what fails, then stats count(*) by bin(5m) to see when it started, and pct(latency_ms, 99) by path for what is slow. From the CLI that is start-query and get-query-results until Complete. The start time is what I correlate with deploys and CloudTrail.
Also asked: How would you find which of two services started failing first? · How do you extract fields from plain-text logs in Logs Insights? · How do you keep Logs Insights queries cheap?
Learn it: 31.11 Logs Insights: query your logs like a database
Someone deleted a production IAM role. How do you find out who? Mid
CloudTrail lookup-events for EventName DeleteRole with --region us-east-1, because IAM is global. The record's userIdentity gives the principal - for an assumed role the session issuer and the session name, which with Identity Center is the person - plus the source IP and the user agent. The Detach and DeleteRolePolicy events just before show what the role had; beyond 90 days the trail's files or log group have the CreateRole.
Also asked: How would you notice that someone stopped your CloudTrail trail? · What are CloudTrail data events, and when do you need them? · Why would you send CloudTrail to a CloudWatch Logs group?
Learn it: 31.14 CloudTrail: event history, trails and reading a record
The AWS bill doubled this month. How do you find out why? Mid
Cost Explorer get-cost-and-usage grouped by SERVICE, DAILY, to find the line and the day it jumped; then that service grouped by USAGE_TYPE - often it is EC2 - Other, which is NAT gateway bytes, EBS or data transfer; then by tag or resource to find the owner. The day usually matches a deploy or a new job. Typical causes: data through a NAT gateway, cross-AZ traffic, forgotten resources, log ingestion.
Also asked: Which data transfer on AWS is free, and which is not? · How would you cut a NAT gateway bill? · What costs money on AWS even when nothing uses it?
Learn it: 31.18 The cost model: what you pay for, and the bills that surprise people
How do you make teams own their AWS costs? Mid
Tag every resource with a team and an environment, activate those as cost allocation tags (update-cost-allocation-tags-status) and backfill, keep the untagged share small with a tag policy for the spelling and IaC rules for the presence, give each team a monthly budget filtered on its tag with an actual and a forecasted alert, and review the cost by team in Cost Explorer every month.
Also asked: What does an AWS tag policy enforce, and what does it not? · Why might a budget never alert even though the team overspends? · What is cost anomaly detection good for?
Learn it: 31.22 Tags, cost allocation and budgets
What is a service control policy, and how is it different from an IAM policy? Mid
An SCP is an Organizations guardrail attached to the root, an OU or an account: it never grants, it limits the maximum anyone in those accounts can do, the root user included, except in the management account. An action needs an allow from an SCP at every level and no deny anywhere, and then still an IAM allow. A typical one denies everything outside approved Regions, with NotAction exemptions for global services like IAM.
Also asked: Why does an AWS organization have many accounts instead of one? · What happens if you detach FullAWSAccess from an OU? · How would you stop anyone from disabling CloudTrail in every account?
The root volume of an EC2 instance is full and you cannot reboot it. What do you do? Mid
Free a little space if something is safe to delete, then aws ec2 modify-volume to a bigger size; once the modification is optimizing the instance sees the bigger disk, so sudo growpart /dev/nvme0n1 1 and sudo resize2fs on the partition (or xfs_growfs for XFS), online. Then fix why it filled up - rotation, retention - and add a CloudWatch agent disk alarm with lead time.
Also asked: Why does EC2 not publish disk usage, and how do you monitor it? · What changed between gp2 and gp3 volumes? · When would you take a snapshot before changing a volume?
Learn it: 31.29 EBS from the instance: gp3, growing a disk online, and the disk metrics EC2 does not have
How would you make a web service on AWS highly available? Mid
Run it in at least two Availability Zones behind a load balancer with health checks, in an Auto Scaling group sized so the surviving zones carry the load if one fails (static stability), with managed data stores in Multi-AZ mode and backups copied to another Region or account and restored regularly. Then alarms on user-facing symptoms, agreed RTO and RPO targets, and game days to prove it.
Also asked: What is the difference between RTO and RPO? · How do you tell whether an outage is yours or AWS's? · Walk me through how you would work an incident on AWS.
Learn it: 31.32 Well-Architected reliability: multi-AZ, backups and the incident on AWS
Practise these answers with flashcards and labs Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.