CloudWatch metrics: namespaces, dimensions, statistics
It is 09:40 and someone in the channel writes "checkout feels slow". Before you read a single log line you want numbers: how many requests, how many failed, how long they took, and since when. On AWS those numbers are in Amazon CloudWatch, the metric store every service writes into. This lesson is about asking CloudWatch the right question and reading the answer correctly - which is harder than it looks, because the API is full of quiet traps.
Need to know: a metric is a namespace (AWS/ApplicationELB), a name (HTTPCode_Target_5XX_Count) and a set of dimensions (LoadBalancer=app/shop-alb/...), and the dimensions must match exactly. You ask for a statistic (Sum, Average, Maximum, SampleCount, p99) per period (60 s, 300 s ...). get-metric-statistics returns its datapoints in no particular order and at most 1,440 of them. EC2 publishes every 5 minutes by default and has no memory or disk metrics at all. Count metrics such as an ALB's 5xx are only published when they are non-zero, so "no data" usually means "zero".
What a metric is
A metric is a time series: (time, value) pairs. CloudWatch names one with three parts:
- the namespace - who publishes it:
AWS/EC2,AWS/ApplicationELB,AWS/NATGateway,AWS/Lambda; your own (custom) metrics use any name that does not start withAWS/; - the metric name -
CPUUtilization,RequestCount,TargetResponseTime; - the dimensions - name=value pairs that say which one:
InstanceId=i-0a1b...,LoadBalancer=app/shop-alb/50dc6c495c0c9188. Every distinct combination of dimensions is a separate metric.
list-metrics shows what exists (metrics that received data in the last two weeks). It never shows values:
$ aws cloudwatch list-metrics --namespace AWS/ApplicationELB --query 'Metrics[].[MetricName,Dimensions[0].Name,Dimensions[0].Value]' --output text | sort
HTTPCode_ELB_5XX_Count LoadBalancer app/shop-alb/50dc6c495c0c9188
HTTPCode_Target_4XX_Count LoadBalancer app/shop-alb/50dc6c495c0c9188
HTTPCode_Target_5XX_Count LoadBalancer app/shop-alb/50dc6c495c0c9188
HealthyHostCount TargetGroup targetgroup/shop-web/6d0ecf831eec9f09
RequestCount LoadBalancer app/shop-alb/50dc6c495c0c9188
TargetResponseTime LoadBalancer app/shop-alb/50dc6c495c0c9188
UnHealthyHostCount TargetGroup targetgroup/shop-web/6d0ecf831eec9f09
$ aws cloudwatch list-metrics --metric-name CPUUtilization --query 'Metrics[].Dimensions[].Value' --output text
i-0a1b2c3d4e5f60718 i-0b2c3d4e5f6071829
The shop has one Application Load Balancer (app/shop-alb/50dc6c495c0c9188 is the part of its ARN after loadbalancer/) and two web instances behind it. The target group's metrics carry two dimensions (TargetGroup and LoadBalancer): to read them you must pass both.
Asking for numbers
get-metric-statistics reads one metric: a time window, a period and one or more statistics. A time window in the shell is easiest as two variables:
$ S=$(date -u -d '-30 min' +%FT%TZ); E=$(date -u +%FT%TZ); echo "$S -> $E"
2026-09-22T19:30:04Z -> 2026-09-22T20:00:04Z
$ aws cloudwatch get-metric-statistics --namespace AWS/ApplicationELB --metric-name RequestCount --dimensions Name=LoadBalancer,Value=app/shop-alb/50dc6c495c0c9188 --start-time $S --end-time $E --period 300 --statistics Sum --output table
--------------------------------------------------
| GetMetricStatistics |
+----------------+-------------------------------+
| Label | RequestCount |
+----------------+-------------------------------+
|| Datapoints ||
|+------+-----------------------------+---------+|
|| Sum | Timestamp | Unit ||
|+------+-----------------------------+---------+|
|| 7422| 2026-09-22T19:35:00+00:00 | Count ||
|| 7401| 2026-09-22T19:45:00+00:00 | Count ||
|| 7343| 2026-09-22T19:55:00+00:00 | Count ||
|| 7386| 2026-09-22T19:50:00+00:00 | Count ||
|| 7558| 2026-09-22T19:30:00+00:00 | Count ||
|| 7492| 2026-09-22T19:40:00+00:00 | Count ||
|| 1411| 2026-09-22T20:00:00+00:00 | Count ||
|+------+-----------------------------+---------+|
Read the timestamps: they are not in order. CloudWatch returns Datapoints unordered, and it is up to you to sort them - with --query 'sort_by(Datapoints, &Timestamp)'. A script that takes Datapoints[-1] as "the latest value" without sorting reads a random period. The last period is also partial: the five minutes are not over yet, so its Sum is small and will grow.
$ aws cloudwatch get-metric-statistics --namespace AWS/ApplicationELB --metric-name RequestCount --dimensions Name=LoadBalancer,Value=app/shop-alb/50dc6c495c0c9188 --start-time $S --end-time $E --period 300 --statistics Sum --query 'sort_by(Datapoints, &Timestamp)[].[Timestamp,Sum]' --output text
2026-09-22T19:30:00+00:00 7558
2026-09-22T19:35:00+00:00 7422
2026-09-22T19:40:00+00:00 7492
2026-09-22T19:45:00+00:00 7401
2026-09-22T19:50:00+00:00 7386
2026-09-22T19:55:00+00:00 7343
2026-09-22T20:00:00+00:00 1411
The statistics, and when each one is the right one:
| statistic | what it is | use it for |
|---|---|---|
Sum | all values in the period added up | counts: requests, errors, bytes |
Average | Sum / SampleCount | utilisation (CPU), but it hides spikes |
Maximum / Minimum | the extremes | "did it ever hit 100%?" |
SampleCount | how many samples were published | sanity checks |
p50, p90, p99 (--extended-statistics) | percentiles | latency: what the slow requests saw |
Latency is where Average lies the most. The ALB's TargetResponseTime (seconds) has a nice average and an ugly tail:
$ aws cloudwatch get-metric-statistics --namespace AWS/ApplicationELB --metric-name TargetResponseTime --dimensions Name=LoadBalancer,Value=app/shop-alb/50dc6c495c0c9188 --start-time $S --end-time $E --period 300 --statistics Average --extended-statistics p99 --query 'sort_by(Datapoints, &Timestamp)[].[Timestamp,Average,ExtendedStatistics.p99]' --output table
-----------------------------------------------------
| GetMetricStatistics |
+----------------------------+----------+-----------+
| 2026-09-22T19:30:00+00:00 | 0.0628 | 0.25746 |
| 2026-09-22T19:35:00+00:00 | 0.05992 | 0.24566 |
| 2026-09-22T19:40:00+00:00 | 0.05836 | 0.23926 |
| 2026-09-22T19:45:00+00:00 | 0.06112 | 0.25058 |
| 2026-09-22T19:50:00+00:00 | 0.06232 | 0.25552 |
| 2026-09-22T19:55:00+00:00 | 0.06054 | 0.24822 |
| 2026-09-22T20:00:00+00:00 | 0.0562 | 0.2304 |
+----------------------------+----------+-----------+
An average of 60 ms and a p99 of 250 ms means one request in a hundred waits four times longer than "typical". Users notice the p99; alarms on latency should use it.
Periods, resolution and the 1,440 limit
The period is the bucket size, in seconds, a multiple of 60 for normal metrics. You choose it per request, but the data underneath has a fixed resolution and a retention:
- data points at 1 minute are kept 15 days, then aggregated to 5 minutes;
- 5 minutes are kept 63 days, then 1 hour;
- 1 hour is kept 455 days (15 months);
- high-resolution custom metrics (1 second) are kept 3 hours.
So "show me last month per minute" is not possible, and asking for too many points at once is refused:
$ aws cloudwatch get-metric-statistics --namespace AWS/ApplicationELB --metric-name RequestCount --dimensions Name=LoadBalancer,Value=app/shop-alb/50dc6c495c0c9188 --start-time $(date -u -d '-2 days' +%FT%TZ) --end-time $E --period 60 --statistics Sum
aws: [ERROR]: An error occurred (InvalidParameterCombination) when calling the GetMetricStatistics operation: You have requested up to 2,880 datapoints, which exceeds the limit of 1,440. You may reduce the datapoints requested by increasing Period, or decreasing the time range.
Two days at one minute is 2,880 points; one call returns at most 1,440. Use a bigger period (300 over two days) or get-metric-data, which pages through up to 100,800 points.
EC2's own metrics come every 5 minutes (basic monitoring, free); detailed monitoring (1 minute) costs extra per instance. Ask for --period 60 on a basic-monitoring instance and you get one point in five minutes. And EC2 sees the instance from the outside: CPU, network, disk I/O operations, status checks - nothing about memory or how full the disks are. Those come from the CloudWatch agent on the instance, which publishes mem_used_percent and disk_used_percent into the CWAgent namespace (lesson 9 uses it).
Dimensions must match exactly
The trap that costs the most time: a metric published with dimensions A and B is a different metric from one with only A. Ask with the wrong set and CloudWatch does not complain - it returns nothing:
$ aws cloudwatch get-metric-statistics --namespace AWS/ApplicationELB --metric-name HealthyHostCount --dimensions Name=LoadBalancer,Value=app/shop-alb/50dc6c495c0c9188 --start-time $S --end-time $E --period 300 --statistics Minimum
{
"Label": "HealthyHostCount",
"Datapoints": []
}
$ aws cloudwatch get-metric-statistics --namespace AWS/ApplicationELB --metric-name HealthyHostCount --dimensions Name=TargetGroup,Value=targetgroup/shop-web/6d0ecf831eec9f09 Name=LoadBalancer,Value=app/shop-alb/50dc6c495c0c9188 --start-time $S --end-time $E --period 300 --statistics Minimum --query 'Datapoints[0].Minimum'
2
Empty Datapoints means "no metric with exactly these dimensions had data", not "the value is zero". When in doubt, list-metrics --metric-name X shows the dimension sets that exist.
Metric math: rates, not counts
"52 errors in five minutes" means little without the traffic. get-metric-data reads several metrics in one call and computes expressions over them. Each query has an Id; queries with ReturnData: false feed the expression without being printed. The error rate in percent:
$ cat ~/oncall-lab/labs/aws/try3/error-rate.json
[
{
"Id": "errors",
"MetricStat": {
"Metric": {
"Namespace": "AWS/ApplicationELB",
"MetricName": "HTTPCode_Target_5XX_Count",
"Dimensions": [
{
"Name": "LoadBalancer",
"Value": "app/shop-alb/50dc6c495c0c9188"
}
]
},
"Period": 300,
"Stat": "Sum"
},
"ReturnData": false
},
{
"Id": "requests",
"MetricStat": {
"Metric": {
"Namespace": "AWS/ApplicationELB",
"MetricName": "RequestCount",
"Dimensions": [
{
"Name": "LoadBalancer",
"Value": "app/shop-alb/50dc6c495c0c9188"
}
]
},
"Period": 300,
"Stat": "Sum"
},
"ReturnData": false
},
{
"Id": "rate",
"Expression": "100 * FILL(errors, 0) / requests",
"Label": "5xx %"
}
]
$ aws cloudwatch get-metric-data --metric-data-queries file://$HOME/oncall-lab/labs/aws/try3/error-rate.json --start-time $S --end-time $E --query 'MetricDataResults[0].{label: Label, newest: Timestamps[:3], percent: Values[:3]}'
{
"label": "5xx %",
"newest": [
"2026-09-22T20:00:00+00:00",
"2026-09-22T19:55:00+00:00",
"2026-09-22T19:50:00+00:00"
],
"percent": [
0,
0.0136184121,
0.0135391281
]
}
get-metric-data returns its timestamps sorted - newest first - which is why [0] is the latest. The FILL(errors, 0) matters because of the next gotcha.
Sparse metrics: no data means zero
The ALB's HTTPCode_Target_5XX_Count is only published in minutes that had at least one 5xx. A healthy minute has no datapoint at all:
$ aws cloudwatch get-metric-statistics --namespace AWS/ApplicationELB --metric-name HTTPCode_Target_5XX_Count --dimensions Name=LoadBalancer,Value=app/shop-alb/50dc6c495c0c9188 --start-time $S --end-time $E --period 60 --statistics Sum --query 'length(Datapoints)'
11
Fewer datapoints than minutes: the minutes without a datapoint had no 5xx. Without FILL, the rate expression has no value in those minutes - and an alarm on it sits in INSUFFICIENT_DATA exactly when everything is fine (the next lesson). FILL(m, 0) turns missing points into zeros; other metric math functions: SUM(METRICS()), RATE(m), IF(cond, a, b).
Your own metrics
Anything you can count, you can publish: queue depth, jobs processed, the age of the last backup. put-metric-data creates the metric on first use:
$ aws cloudwatch put-metric-data --namespace Try/Shop --metric-name OrdersQueued --dimensions Service=checkout --value 42 --unit Count
$ aws cloudwatch put-metric-data --namespace AWS/Shop --metric-name OrdersQueued --value 1
aws: [ERROR]: An error occurred (InvalidParameterValue) when calling the PutMetricData operation: The value AWS/Shop for parameter Namespace is invalid.
Two things to know before you script it: every unique combination of namespace, name and dimensions is a billed metric (about $0.30 a month each), so a dimension like RequestId or UserId turns into thousands of metrics and a surprise bill; and AWS/ is reserved.
In an interview: "Average latency is fine but users complain - what do you look at?" - "The percentiles, p99 or p95 of the load balancer's TargetResponseTime: the average hides the slow tail that users actually feel. Then break it down by target or path to find who is slow."
You can now: find a metric with list-metrics, read it with get-metric-statistics (the right statistic, a sensible period, sorted), explain why a query returns nothing (exact dimensions, sparse counts, basic monitoring), compute a rate with get-metric-data, and publish a custom metric without exploding the bill.