OnCallReady

Lesson 31.1 · AWS III: CloudWatch, CloudTrail, Cost & Incidents · 22 min read

CloudWatch metrics: namespaces, dimensions, statistics and periods

In plain words

A metric is a gauge that writes down its needle position every minute: how many requests, how much CPU, how slow. CloudWatch keeps all of those notes for every service in the account.

To read one, you say which gauge (its namespace and name), which copy of it (the dimensions, like which load balancer), how to summarise each stretch of time (a sum, an average, the slowest one in a hundred), and how long each stretch is. Ask with the wrong label and CloudWatch quietly shows you nothing at all.

CloudWatch metrics: namespaces, dimensions, statistics

It is 09:40 and someone in the channel writes "checkout feels slow". Before you read a single log line you want numbers: how many requests, how many failed, how long they took, and since when. On AWS those numbers are in Amazon CloudWatch, the metric store every service writes into. This lesson is about asking CloudWatch the right question and reading the answer correctly - which is harder than it looks, because the API is full of quiet traps.

Need to know: a metric is a namespace (AWS/ApplicationELB), a name (HTTPCode_Target_5XX_Count) and a set of dimensions (LoadBalancer=app/shop-alb/...), and the dimensions must match exactly. You ask for a statistic (Sum, Average, Maximum, SampleCount, p99) per period (60 s, 300 s ...). get-metric-statistics returns its datapoints in no particular order and at most 1,440 of them. EC2 publishes every 5 minutes by default and has no memory or disk metrics at all. Count metrics such as an ALB's 5xx are only published when they are non-zero, so "no data" usually means "zero".

What a metric is

A metric is a time series: (time, value) pairs. CloudWatch names one with three parts:

list-metrics shows what exists (metrics that received data in the last two weeks). It never shows values:

$ aws cloudwatch list-metrics --namespace AWS/ApplicationELB --query 'Metrics[].[MetricName,Dimensions[0].Name,Dimensions[0].Value]' --output text | sort
HTTPCode_ELB_5XX_Count	LoadBalancer	app/shop-alb/50dc6c495c0c9188
HTTPCode_Target_4XX_Count	LoadBalancer	app/shop-alb/50dc6c495c0c9188
HTTPCode_Target_5XX_Count	LoadBalancer	app/shop-alb/50dc6c495c0c9188
HealthyHostCount	TargetGroup	targetgroup/shop-web/6d0ecf831eec9f09
RequestCount	LoadBalancer	app/shop-alb/50dc6c495c0c9188
TargetResponseTime	LoadBalancer	app/shop-alb/50dc6c495c0c9188
UnHealthyHostCount	TargetGroup	targetgroup/shop-web/6d0ecf831eec9f09
$ aws cloudwatch list-metrics --metric-name CPUUtilization --query 'Metrics[].Dimensions[].Value' --output text
i-0a1b2c3d4e5f60718	i-0b2c3d4e5f6071829

The shop has one Application Load Balancer (app/shop-alb/50dc6c495c0c9188 is the part of its ARN after loadbalancer/) and two web instances behind it. The target group's metrics carry two dimensions (TargetGroup and LoadBalancer): to read them you must pass both.

Asking for numbers

get-metric-statistics reads one metric: a time window, a period and one or more statistics. A time window in the shell is easiest as two variables:

$ S=$(date -u -d '-30 min' +%FT%TZ); E=$(date -u +%FT%TZ); echo "$S -> $E"
2026-09-22T19:30:04Z -> 2026-09-22T20:00:04Z
$ aws cloudwatch get-metric-statistics --namespace AWS/ApplicationELB --metric-name RequestCount --dimensions Name=LoadBalancer,Value=app/shop-alb/50dc6c495c0c9188 --start-time $S --end-time $E --period 300 --statistics Sum --output table
--------------------------------------------------
|               GetMetricStatistics              |
+----------------+-------------------------------+
|  Label         |  RequestCount                 |
+----------------+-------------------------------+
||                  Datapoints                  ||
|+------+-----------------------------+---------+|
||  Sum |          Timestamp          |  Unit   ||
|+------+-----------------------------+---------+|
||  7422|  2026-09-22T19:35:00+00:00  |  Count  ||
||  7401|  2026-09-22T19:45:00+00:00  |  Count  ||
||  7343|  2026-09-22T19:55:00+00:00  |  Count  ||
||  7386|  2026-09-22T19:50:00+00:00  |  Count  ||
||  7558|  2026-09-22T19:30:00+00:00  |  Count  ||
||  7492|  2026-09-22T19:40:00+00:00  |  Count  ||
||  1411|  2026-09-22T20:00:00+00:00  |  Count  ||
|+------+-----------------------------+---------+|

Read the timestamps: they are not in order. CloudWatch returns Datapoints unordered, and it is up to you to sort them - with --query 'sort_by(Datapoints, &Timestamp)'. A script that takes Datapoints[-1] as "the latest value" without sorting reads a random period. The last period is also partial: the five minutes are not over yet, so its Sum is small and will grow.

$ aws cloudwatch get-metric-statistics --namespace AWS/ApplicationELB --metric-name RequestCount --dimensions Name=LoadBalancer,Value=app/shop-alb/50dc6c495c0c9188 --start-time $S --end-time $E --period 300 --statistics Sum --query 'sort_by(Datapoints, &Timestamp)[].[Timestamp,Sum]' --output text
2026-09-22T19:30:00+00:00	7558
2026-09-22T19:35:00+00:00	7422
2026-09-22T19:40:00+00:00	7492
2026-09-22T19:45:00+00:00	7401
2026-09-22T19:50:00+00:00	7386
2026-09-22T19:55:00+00:00	7343
2026-09-22T20:00:00+00:00	1411

The statistics, and when each one is the right one:

statisticwhat it isuse it for
Sumall values in the period added upcounts: requests, errors, bytes
AverageSum / SampleCountutilisation (CPU), but it hides spikes
Maximum / Minimumthe extremes"did it ever hit 100%?"
SampleCounthow many samples were publishedsanity checks
p50, p90, p99 (--extended-statistics)percentileslatency: what the slow requests saw

Latency is where Average lies the most. The ALB's TargetResponseTime (seconds) has a nice average and an ugly tail:

$ aws cloudwatch get-metric-statistics --namespace AWS/ApplicationELB --metric-name TargetResponseTime --dimensions Name=LoadBalancer,Value=app/shop-alb/50dc6c495c0c9188 --start-time $S --end-time $E --period 300 --statistics Average --extended-statistics p99 --query 'sort_by(Datapoints, &Timestamp)[].[Timestamp,Average,ExtendedStatistics.p99]' --output table
-----------------------------------------------------
|                GetMetricStatistics                |
+----------------------------+----------+-----------+
|  2026-09-22T19:30:00+00:00 |  0.0628  |  0.25746  |
|  2026-09-22T19:35:00+00:00 |  0.05992 |  0.24566  |
|  2026-09-22T19:40:00+00:00 |  0.05836 |  0.23926  |
|  2026-09-22T19:45:00+00:00 |  0.06112 |  0.25058  |
|  2026-09-22T19:50:00+00:00 |  0.06232 |  0.25552  |
|  2026-09-22T19:55:00+00:00 |  0.06054 |  0.24822  |
|  2026-09-22T20:00:00+00:00 |  0.0562  |  0.2304   |
+----------------------------+----------+-----------+

An average of 60 ms and a p99 of 250 ms means one request in a hundred waits four times longer than "typical". Users notice the p99; alarms on latency should use it.

Periods, resolution and the 1,440 limit

The period is the bucket size, in seconds, a multiple of 60 for normal metrics. You choose it per request, but the data underneath has a fixed resolution and a retention:

So "show me last month per minute" is not possible, and asking for too many points at once is refused:

$ aws cloudwatch get-metric-statistics --namespace AWS/ApplicationELB --metric-name RequestCount --dimensions Name=LoadBalancer,Value=app/shop-alb/50dc6c495c0c9188 --start-time $(date -u -d '-2 days' +%FT%TZ) --end-time $E --period 60 --statistics Sum
aws: [ERROR]: An error occurred (InvalidParameterCombination) when calling the GetMetricStatistics operation: You have requested up to 2,880 datapoints, which exceeds the limit of 1,440. You may reduce the datapoints requested by increasing Period, or decreasing the time range.

Two days at one minute is 2,880 points; one call returns at most 1,440. Use a bigger period (300 over two days) or get-metric-data, which pages through up to 100,800 points.

EC2's own metrics come every 5 minutes (basic monitoring, free); detailed monitoring (1 minute) costs extra per instance. Ask for --period 60 on a basic-monitoring instance and you get one point in five minutes. And EC2 sees the instance from the outside: CPU, network, disk I/O operations, status checks - nothing about memory or how full the disks are. Those come from the CloudWatch agent on the instance, which publishes mem_used_percent and disk_used_percent into the CWAgent namespace (lesson 9 uses it).

Dimensions must match exactly

The trap that costs the most time: a metric published with dimensions A and B is a different metric from one with only A. Ask with the wrong set and CloudWatch does not complain - it returns nothing:

$ aws cloudwatch get-metric-statistics --namespace AWS/ApplicationELB --metric-name HealthyHostCount --dimensions Name=LoadBalancer,Value=app/shop-alb/50dc6c495c0c9188 --start-time $S --end-time $E --period 300 --statistics Minimum
{
    "Label": "HealthyHostCount",
    "Datapoints": []
}
$ aws cloudwatch get-metric-statistics --namespace AWS/ApplicationELB --metric-name HealthyHostCount --dimensions Name=TargetGroup,Value=targetgroup/shop-web/6d0ecf831eec9f09 Name=LoadBalancer,Value=app/shop-alb/50dc6c495c0c9188 --start-time $S --end-time $E --period 300 --statistics Minimum --query 'Datapoints[0].Minimum'
2

Empty Datapoints means "no metric with exactly these dimensions had data", not "the value is zero". When in doubt, list-metrics --metric-name X shows the dimension sets that exist.

Metric math: rates, not counts

"52 errors in five minutes" means little without the traffic. get-metric-data reads several metrics in one call and computes expressions over them. Each query has an Id; queries with ReturnData: false feed the expression without being printed. The error rate in percent:

$ cat ~/oncall-lab/labs/aws/try3/error-rate.json
[
  {
    "Id": "errors",
    "MetricStat": {
      "Metric": {
        "Namespace": "AWS/ApplicationELB",
        "MetricName": "HTTPCode_Target_5XX_Count",
        "Dimensions": [
          {
            "Name": "LoadBalancer",
            "Value": "app/shop-alb/50dc6c495c0c9188"
          }
        ]
      },
      "Period": 300,
      "Stat": "Sum"
    },
    "ReturnData": false
  },
  {
    "Id": "requests",
    "MetricStat": {
      "Metric": {
        "Namespace": "AWS/ApplicationELB",
        "MetricName": "RequestCount",
        "Dimensions": [
          {
            "Name": "LoadBalancer",
            "Value": "app/shop-alb/50dc6c495c0c9188"
          }
        ]
      },
      "Period": 300,
      "Stat": "Sum"
    },
    "ReturnData": false
  },
  {
    "Id": "rate",
    "Expression": "100 * FILL(errors, 0) / requests",
    "Label": "5xx %"
  }
]
$ aws cloudwatch get-metric-data --metric-data-queries file://$HOME/oncall-lab/labs/aws/try3/error-rate.json --start-time $S --end-time $E --query 'MetricDataResults[0].{label: Label, newest: Timestamps[:3], percent: Values[:3]}'
{
    "label": "5xx %",
    "newest": [
        "2026-09-22T20:00:00+00:00",
        "2026-09-22T19:55:00+00:00",
        "2026-09-22T19:50:00+00:00"
    ],
    "percent": [
        0,
        0.0136184121,
        0.0135391281
    ]
}

get-metric-data returns its timestamps sorted - newest first - which is why [0] is the latest. The FILL(errors, 0) matters because of the next gotcha.

Sparse metrics: no data means zero

The ALB's HTTPCode_Target_5XX_Count is only published in minutes that had at least one 5xx. A healthy minute has no datapoint at all:

$ aws cloudwatch get-metric-statistics --namespace AWS/ApplicationELB --metric-name HTTPCode_Target_5XX_Count --dimensions Name=LoadBalancer,Value=app/shop-alb/50dc6c495c0c9188 --start-time $S --end-time $E --period 60 --statistics Sum --query 'length(Datapoints)'
11

Fewer datapoints than minutes: the minutes without a datapoint had no 5xx. Without FILL, the rate expression has no value in those minutes - and an alarm on it sits in INSUFFICIENT_DATA exactly when everything is fine (the next lesson). FILL(m, 0) turns missing points into zeros; other metric math functions: SUM(METRICS()), RATE(m), IF(cond, a, b).

Your own metrics

Anything you can count, you can publish: queue depth, jobs processed, the age of the last backup. put-metric-data creates the metric on first use:

$ aws cloudwatch put-metric-data --namespace Try/Shop --metric-name OrdersQueued --dimensions Service=checkout --value 42 --unit Count
$ aws cloudwatch put-metric-data --namespace AWS/Shop --metric-name OrdersQueued --value 1
aws: [ERROR]: An error occurred (InvalidParameterValue) when calling the PutMetricData operation: The value AWS/Shop for parameter Namespace is invalid.

Two things to know before you script it: every unique combination of namespace, name and dimensions is a billed metric (about $0.30 a month each), so a dimension like RequestId or UserId turns into thousands of metrics and a surprise bill; and AWS/ is reserved.

In an interview: "Average latency is fine but users complain - what do you look at?" - "The percentiles, p99 or p95 of the load balancer's TargetResponseTime: the average hides the slow tail that users actually feel. Then break it down by target or path to find who is slow."

You can now: find a metric with list-metrics, read it with get-metric-statistics (the right statistic, a sensible period, sorted), explain why a query returns nothing (exact dimensions, sparse counts, basic monitoring), compute a rate with get-metric-data, and publish a custom metric without exploding the bill.

Why it helps

Every investigation on AWS starts with numbers: is it broken, since when, how badly, for whom. Reading CloudWatch correctly is what turns "checkout feels slow" into "p99 latency went from 250 ms to 3 s at 09:12 on one load balancer".

The traps are quiet ones - unordered datapoints, dimensions that must match exactly, metrics that are only published when they are not zero, five-minute EC2 data - and each of them has misled someone in a real incident.

Commands in this lesson

aws cat

FAQ

Why are the datapoints from get-metric-statistics out of order?

The API returns them unordered; it has always done that. Sort them yourself with --query 'sort_by(Datapoints, &Timestamp)' before you read the latest value, or you may report an old minute as the current one. get-metric-data, the newer call, returns its timestamps sorted, newest first.

Why does my query return no datapoints when the service is clearly busy?

Most often the dimensions do not match exactly - the metric was published with two dimensions and you asked with one, or none. Run list-metrics with the metric name to see the dimension sets that exist. Other causes are a sparse metric (no 5xx in that minute), basic five-minute monitoring, or a --unit that does not match.

Which statistic should I use?

Sum for counts such as requests, errors and bytes; Average or Maximum for utilisation such as CPU; a percentile (p99, p95) for latency. Averages hide the slow tail that users feel, so latency alarms and dashboards should use percentiles.

Where are memory and disk usage for my EC2 instance?

EC2 does not publish them: it measures the instance from the hypervisor, so it knows CPU, network, disk operations and status checks, but not what is inside the filesystem or memory. The CloudWatch agent on the instance publishes mem_used_percent and disk_used_percent into the CWAgent namespace.

Do custom metrics cost money?

Yes, roughly $0.30 per metric per month, and every distinct combination of namespace, name and dimensions is one metric. A dimension with many values, such as a user or request ID, turns into thousands of metrics. Keep dimensions to a handful of values.

In an interview Mid

Average latency looks fine but users complain about slowness. What do you look at?

The percentiles: p99 or p95 of the load balancer's TargetResponseTime with get-metric-statistics --extended-statistics p99. An average of 60 ms can hide a p99 of several seconds - one request in a hundred is slow, and those are the users who complain. Then break it down by target or path to find who is slow, and alarm on the percentile rather than the average.

Also asked: Why might a CloudWatch query return no data for a metric you know exists? · How do you compute an error rate from two CloudWatch metrics? · What metrics does EC2 not publish, and how do you get them?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.