Alarms: states, M out of N, missing data and the page
A metric nobody looks at is a graph in a dashboard nobody opens. A CloudWatch alarm watches one metric (or an expression) for you and changes state when it crosses a threshold; a state change runs actions - most often a message to an SNS topic that emails or pages the on-call engineer. This lesson builds alarms on the shop's load balancer and makes them page, and shows the two settings behind most "the alarm never fired" post-mortems.
Need to know: an alarm is in one of three states - OK, ALARM, INSUFFICIENT_DATA (a new alarm starts here). Every minute it looks at the last N periods (--evaluation-periods) and goes to ALARM when M of them breach (--datapoints-to-alarm). --treat-missing-data decides what a period without data means: missing (the default: no decision, possibly INSUFFICIENT_DATA), notBreaching, breaching or ignore. Actions run on a state change only, once. An email subscription to an SNS topic delivers nothing until someone confirms it.
The anatomy of an alarm
put-metric-alarm names the metric exactly like get-metric-statistics does, then says what is bad:
| option | meaning |
|---|---|
--namespace, --metric-name, --dimensions | which metric (exact dimensions again) |
--statistic Sum (or --extended-statistic p99) | how to aggregate each period |
--period 60 | the period length in seconds |
--evaluation-periods 3 | N: how many recent periods to look at |
--datapoints-to-alarm 2 | M: how many of those must breach (default: N) |
--threshold 5 --comparison-operator GreaterThanThreshold | what breaching means |
--treat-missing-data notBreaching | what an empty period means |
--alarm-actions arn:aws:sns:... | what to do on the way into ALARM (also --ok-actions) |
"2 out of 3 one-minute periods with more than 5 errors" ignores a single noisy minute but fires within three minutes of a real problem. That M-out-of-N shape is the usual answer to flapping alarms: not a higher threshold, a longer look.
The alarm that never fires
The shop's 5xx count had a burst twenty minutes ago and has been quiet since. An alarm with the defaults:
$ aws cloudwatch put-metric-alarm --alarm-name try-5xx --namespace AWS/ApplicationELB --metric-name HTTPCode_Target_5XX_Count --dimensions Name=LoadBalancer,Value=app/shop-alb/50dc6c495c0c9188 --statistic Sum --period 60 --evaluation-periods 3 --datapoints-to-alarm 2 --threshold 5 --comparison-operator GreaterThanThreshold
$ aws cloudwatch describe-alarms --alarm-names try-5xx --query 'MetricAlarms[0].[StateValue,StateReason]' --output text
INSUFFICIENT_DATA Unchecked: Initial alarm creation
$ sleep 70
$ aws cloudwatch describe-alarms --alarm-names try-5xx --query 'MetricAlarms[0].[StateValue,StateReason]' --output text
INSUFFICIENT_DATA Unchecked: Initial alarm creation
A new alarm is INSUFFICIENT_DATA with "Unchecked: Initial alarm creation". After its first evaluation it is still INSUFFICIENT_DATA: the 5xx metric is only published in minutes with a 5xx (the previous lesson), the last three minutes had none, so there is no data - and missing means "no data, no decision". Many teams route INSUFFICIENT_DATA nowhere, so this alarm is silent when things are fine and it would take the first breaching minutes to wake up. For a sparse count, missing data is good news:
$ aws cloudwatch put-metric-alarm --alarm-name try-5xx --namespace AWS/ApplicationELB --metric-name HTTPCode_Target_5XX_Count --dimensions Name=LoadBalancer,Value=app/shop-alb/50dc6c495c0c9188 --statistic Sum --period 60 --evaluation-periods 3 --datapoints-to-alarm 2 --threshold 5 --comparison-operator GreaterThanThreshold --treat-missing-data notBreaching
$ sleep 60
$ aws cloudwatch describe-alarms --alarm-names try-5xx --query 'MetricAlarms[0].[StateValue,StateReason]' --output text
OK Threshold Crossed: no datapoints were received for 3 periods and 3 missing datapoints were treated as [NonBreaching].
The four choices, and when each is right:
notBreaching- errors, 5xx, failed jobs: no datapoint means nothing went wrong;breaching- heartbeats and "the job ran" metrics: silence is the problem (a backup that published nothing did not run);ignore- keep the current state until data comes back (rarely what you want);missing- the default; fine for metrics that are always published (CPU, RequestCount), where a gap really means "we do not know".
Running put-metric-alarm again with the same name updates the alarm (the history says "updated"); there is no separate update command.
Reading the state
StateReason is the sentence to read first; StateReasonData has the numbers behind it. The history keeps 30 days of state changes, configuration updates and actions:
$ aws cloudwatch describe-alarm-history --alarm-name try-5xx --query 'AlarmHistoryItems[].[Timestamp,HistoryItemType,HistorySummary]' --output text
2026-09-22T20:02:07.000000+00:00 StateUpdate Alarm updated from INSUFFICIENT_DATA to OK
2026-09-22T20:01:14.300000+00:00 ConfigurationUpdate Alarm "try-5xx" updated
2026-09-22T20:00:04.000000+00:00 ConfigurationUpdate Alarm "try-5xx" created
$ aws cloudwatch describe-alarms --alarm-names try-5xx --query 'MetricAlarms[0].StateReasonData' --output text | jq '{statistic, period, threshold, evaluatedDatapoints}'
{
"statistic": "Sum",
"period": 60,
"threshold": 5,
"evaluatedDatapoints": [
{
"timestamp": "2026-09-22T20:00:00.000+0000"
},
{
"timestamp": "2026-09-22T19:59:00.000+0000"
},
{
"timestamp": "2026-09-22T19:58:00.000+0000"
}
]
}
When an alarm "fired late" or "never fired", the history plus StateReasonData answers it in a minute: which datapoints it saw, which were missing, what it decided.
Alarm on the rate, not the count
Five errors in a minute is an outage at night and noise at the Black Friday peak. An alarm can watch a metric math expression: --metrics takes the same queries as get-metric-data, with exactly one of them returning data:
$ cat ~/oncall-lab/labs/aws/try3/rate.json
[
{
"Id": "e",
"MetricStat": {
"Metric": {
"Namespace": "AWS/ApplicationELB",
"MetricName": "HTTPCode_Target_5XX_Count",
"Dimensions": [
{
"Name": "LoadBalancer",
"Value": "app/shop-alb/50dc6c495c0c9188"
}
]
},
"Period": 60,
"Stat": "Sum"
},
"ReturnData": false
},
{
"Id": "r",
"MetricStat": {
"Metric": {
"Namespace": "AWS/ApplicationELB",
"MetricName": "RequestCount",
"Dimensions": [
{
"Name": "LoadBalancer",
"Value": "app/shop-alb/50dc6c495c0c9188"
}
]
},
"Period": 60,
"Stat": "Sum"
},
"ReturnData": false
},
{
"Id": "rate",
"Expression": "100 * FILL(e, 0) / r",
"Label": "5xx percent",
"ReturnData": true
}
]
$ aws cloudwatch put-metric-alarm --alarm-name try-5xx-rate --metrics file://$HOME/oncall-lab/labs/aws/try3/rate.json --evaluation-periods 5 --datapoints-to-alarm 3 --threshold 1 --comparison-operator GreaterThanThreshold --treat-missing-data notBreaching
$ aws cloudwatch describe-alarms --alarm-names try-5xx-rate --query 'MetricAlarms[0].[AlarmName,StateValue,Metrics[2].Expression]' --output text
try-5xx-rate INSUFFICIENT_DATA 100 * FILL(e, 0) / r
Alarm on symptoms users feel - error rate, p99 latency, the queue's oldest message - rather than on causes such as CPU. A CPU alarm at 80% pages you for a busy but healthy box and stays silent for a deadlock at 2% CPU.
Making it page: SNS
An action is an ARN. For notifications it is an SNS topic, and the topic delivers to its subscriptions: email, SMS, an HTTPS endpoint (PagerDuty, Opsgenie), SQS, Lambda. The shop's on-call topic already exists:
$ aws sns list-subscriptions-by-topic --topic-arn arn:aws:sns:eu-central-1:111122223333:oncall-pages --query 'Subscriptions[].[Protocol,Endpoint,SubscriptionArn]' --output text
email [email protected] arn:aws:sns:eu-central-1:111122223333:oncall-pages:28ab2bd4-7eb1-4fa9-abaf-66102563eb6c
$ aws sns create-topic --name try-pages
{
"TopicArn": "arn:aws:sns:eu-central-1:111122223333:try-pages"
}
$ aws sns subscribe --topic-arn arn:aws:sns:eu-central-1:111122223333:try-pages --protocol email --notification-endpoint [email protected]
{
"SubscriptionArn": "pending confirmation"
}
$ aws sns list-subscriptions-by-topic --topic-arn arn:aws:sns:eu-central-1:111122223333:try-pages --query 'Subscriptions[].SubscriptionArn' --output text
PendingConfirmation
PendingConfirmation: SNS sent a confirmation email, and until someone clicks its link the subscription receives nothing - the classic "we set up the alarm and the topic, the page never came". The lab's inbox (simulator) is ~/oncall-lab/labs/aws/pager.log; the confirmation is in it, with the token the link carries:
$ tail -4 ~/oncall-lab/labs/aws/pager.log
Confirm subscription
(simulator) the link carries this token - confirm from the CLI with:
aws sns confirm-subscription --topic-arn arn:aws:sns:eu-central-1:111122223333:try-pages --token 439fd53316671291aa2ada1e7a37495a9c02361be6f3563cf18e747f881a703b
$ aws sns confirm-subscription --topic-arn arn:aws:sns:eu-central-1:111122223333:try-pages --token $(grep -o -- '--token [0-9a-f]*' ~/oncall-lab/labs/aws/pager.log | tail -1 | cut -d' ' -f2)
{
"SubscriptionArn": "arn:aws:sns:eu-central-1:111122223333:try-pages:62a2486b-a695-4c19-b269-99376e74a85a"
}
Now wire the alarm to it and test the page without breaking anything: set-alarm-state puts an alarm into a state by hand, the action runs, and the next evaluation (within a minute) puts the real state back:
$ aws cloudwatch put-metric-alarm --alarm-name try-5xx --namespace AWS/ApplicationELB --metric-name HTTPCode_Target_5XX_Count --dimensions Name=LoadBalancer,Value=app/shop-alb/50dc6c495c0c9188 --statistic Sum --period 60 --evaluation-periods 3 --datapoints-to-alarm 2 --threshold 5 --comparison-operator GreaterThanThreshold --treat-missing-data notBreaching --alarm-actions arn:aws:sns:eu-central-1:111122223333:try-pages --ok-actions arn:aws:sns:eu-central-1:111122223333:try-pages
$ aws cloudwatch set-alarm-state --alarm-name try-5xx --state-value ALARM --state-reason "testing the page (learner)"
$ grep '^Subject' ~/oncall-lab/labs/aws/pager.log | tail -1
Subject: ALARM: "try-5xx" in EU (Frankfurt)
$ sleep 60
$ aws cloudwatch describe-alarm-history --alarm-name try-5xx --history-item-type StateUpdate --max-items 2 --query 'AlarmHistoryItems[].HistorySummary' --output text
Alarm updated from ALARM to OK Alarm updated from OK to ALARM
Three more things you will meet:
- No reminders. An alarm notifies when it enters ALARM. If the page is missed, it does not repeat while the alarm stays in ALARM; the paging tool (PagerDuty, Opsgenie) does the escalation.
- Maintenance:
disable-alarm-actionsstops the pages and keeps the evaluation (the state still changes);enable-alarm-actionsafter the window. Deleting alarms to silence them loses them. - Composite alarms combine alarms with
AND/OR/NOT("page only if error rate AND latency alarm"), the tool for alert fatigue on bigger systems. A standard alarm costs about $0.10 a month per metric it watches.
In an interview: "An alarm never fired during an outage. What do you check?" - "Its history and StateReasonData: was it in INSUFFICIENT_DATA because missing data was left as 'missing' on a sparse metric, were the dimensions exactly the metric's, was M-out-of-N too strict for the period, and did the action work - the SNS topic exists and the subscription is confirmed. Then test with set-alarm-state."
You can now: create an alarm with the right period, M out of N and missing-data treatment, read its state, reason and history, alarm on a rate with metric math, wire it to an SNS topic with a confirmed subscription, and test the page with set-alarm-state.