OnCallReady

Lesson 31.4 · AWS III: CloudWatch, CloudTrail, Cost & Incidents · 19 min read

Alarms: states, M out of N, missing data and notifications

In plain words

An alarm is a smoke detector for one gauge. Every minute it looks at the last few readings and decides: all fine (OK), on fire (ALARM), or "I cannot see anything" (INSUFFICIENT_DATA). You choose how many bad readings out of the last few it takes before it rings, so one puff of smoke does not wake everybody up.

When it changes its mind, it sends a message to a mailbox called an SNS topic, and the topic forwards it to people. A mailbox nobody confirmed delivers nothing.

Alarms: states, M out of N, missing data and the page

A metric nobody looks at is a graph in a dashboard nobody opens. A CloudWatch alarm watches one metric (or an expression) for you and changes state when it crosses a threshold; a state change runs actions - most often a message to an SNS topic that emails or pages the on-call engineer. This lesson builds alarms on the shop's load balancer and makes them page, and shows the two settings behind most "the alarm never fired" post-mortems.

Need to know: an alarm is in one of three states - OK, ALARM, INSUFFICIENT_DATA (a new alarm starts here). Every minute it looks at the last N periods (--evaluation-periods) and goes to ALARM when M of them breach (--datapoints-to-alarm). --treat-missing-data decides what a period without data means: missing (the default: no decision, possibly INSUFFICIENT_DATA), notBreaching, breaching or ignore. Actions run on a state change only, once. An email subscription to an SNS topic delivers nothing until someone confirms it.

The anatomy of an alarm

put-metric-alarm names the metric exactly like get-metric-statistics does, then says what is bad:

optionmeaning
--namespace, --metric-name, --dimensionswhich metric (exact dimensions again)
--statistic Sum (or --extended-statistic p99)how to aggregate each period
--period 60the period length in seconds
--evaluation-periods 3N: how many recent periods to look at
--datapoints-to-alarm 2M: how many of those must breach (default: N)
--threshold 5 --comparison-operator GreaterThanThresholdwhat breaching means
--treat-missing-data notBreachingwhat an empty period means
--alarm-actions arn:aws:sns:...what to do on the way into ALARM (also --ok-actions)

"2 out of 3 one-minute periods with more than 5 errors" ignores a single noisy minute but fires within three minutes of a real problem. That M-out-of-N shape is the usual answer to flapping alarms: not a higher threshold, a longer look.

The alarm that never fires

The shop's 5xx count had a burst twenty minutes ago and has been quiet since. An alarm with the defaults:

$ aws cloudwatch put-metric-alarm --alarm-name try-5xx --namespace AWS/ApplicationELB --metric-name HTTPCode_Target_5XX_Count --dimensions Name=LoadBalancer,Value=app/shop-alb/50dc6c495c0c9188 --statistic Sum --period 60 --evaluation-periods 3 --datapoints-to-alarm 2 --threshold 5 --comparison-operator GreaterThanThreshold
$ aws cloudwatch describe-alarms --alarm-names try-5xx --query 'MetricAlarms[0].[StateValue,StateReason]' --output text
INSUFFICIENT_DATA	Unchecked: Initial alarm creation
$ sleep 70
$ aws cloudwatch describe-alarms --alarm-names try-5xx --query 'MetricAlarms[0].[StateValue,StateReason]' --output text
INSUFFICIENT_DATA	Unchecked: Initial alarm creation

A new alarm is INSUFFICIENT_DATA with "Unchecked: Initial alarm creation". After its first evaluation it is still INSUFFICIENT_DATA: the 5xx metric is only published in minutes with a 5xx (the previous lesson), the last three minutes had none, so there is no data - and missing means "no data, no decision". Many teams route INSUFFICIENT_DATA nowhere, so this alarm is silent when things are fine and it would take the first breaching minutes to wake up. For a sparse count, missing data is good news:

$ aws cloudwatch put-metric-alarm --alarm-name try-5xx --namespace AWS/ApplicationELB --metric-name HTTPCode_Target_5XX_Count --dimensions Name=LoadBalancer,Value=app/shop-alb/50dc6c495c0c9188 --statistic Sum --period 60 --evaluation-periods 3 --datapoints-to-alarm 2 --threshold 5 --comparison-operator GreaterThanThreshold --treat-missing-data notBreaching
$ sleep 60
$ aws cloudwatch describe-alarms --alarm-names try-5xx --query 'MetricAlarms[0].[StateValue,StateReason]' --output text
OK	Threshold Crossed: no datapoints were received for 3 periods and 3 missing datapoints were treated as [NonBreaching].

The four choices, and when each is right:

Running put-metric-alarm again with the same name updates the alarm (the history says "updated"); there is no separate update command.

Reading the state

StateReason is the sentence to read first; StateReasonData has the numbers behind it. The history keeps 30 days of state changes, configuration updates and actions:

$ aws cloudwatch describe-alarm-history --alarm-name try-5xx --query 'AlarmHistoryItems[].[Timestamp,HistoryItemType,HistorySummary]' --output text
2026-09-22T20:02:07.000000+00:00	StateUpdate	Alarm updated from INSUFFICIENT_DATA to OK
2026-09-22T20:01:14.300000+00:00	ConfigurationUpdate	Alarm "try-5xx" updated
2026-09-22T20:00:04.000000+00:00	ConfigurationUpdate	Alarm "try-5xx" created
$ aws cloudwatch describe-alarms --alarm-names try-5xx --query 'MetricAlarms[0].StateReasonData' --output text | jq '{statistic, period, threshold, evaluatedDatapoints}'
{
  "statistic": "Sum",
  "period": 60,
  "threshold": 5,
  "evaluatedDatapoints": [
    {
      "timestamp": "2026-09-22T20:00:00.000+0000"
    },
    {
      "timestamp": "2026-09-22T19:59:00.000+0000"
    },
    {
      "timestamp": "2026-09-22T19:58:00.000+0000"
    }
  ]
}

When an alarm "fired late" or "never fired", the history plus StateReasonData answers it in a minute: which datapoints it saw, which were missing, what it decided.

Alarm on the rate, not the count

Five errors in a minute is an outage at night and noise at the Black Friday peak. An alarm can watch a metric math expression: --metrics takes the same queries as get-metric-data, with exactly one of them returning data:

$ cat ~/oncall-lab/labs/aws/try3/rate.json
[
  {
    "Id": "e",
    "MetricStat": {
      "Metric": {
        "Namespace": "AWS/ApplicationELB",
        "MetricName": "HTTPCode_Target_5XX_Count",
        "Dimensions": [
          {
            "Name": "LoadBalancer",
            "Value": "app/shop-alb/50dc6c495c0c9188"
          }
        ]
      },
      "Period": 60,
      "Stat": "Sum"
    },
    "ReturnData": false
  },
  {
    "Id": "r",
    "MetricStat": {
      "Metric": {
        "Namespace": "AWS/ApplicationELB",
        "MetricName": "RequestCount",
        "Dimensions": [
          {
            "Name": "LoadBalancer",
            "Value": "app/shop-alb/50dc6c495c0c9188"
          }
        ]
      },
      "Period": 60,
      "Stat": "Sum"
    },
    "ReturnData": false
  },
  {
    "Id": "rate",
    "Expression": "100 * FILL(e, 0) / r",
    "Label": "5xx percent",
    "ReturnData": true
  }
]
$ aws cloudwatch put-metric-alarm --alarm-name try-5xx-rate --metrics file://$HOME/oncall-lab/labs/aws/try3/rate.json --evaluation-periods 5 --datapoints-to-alarm 3 --threshold 1 --comparison-operator GreaterThanThreshold --treat-missing-data notBreaching
$ aws cloudwatch describe-alarms --alarm-names try-5xx-rate --query 'MetricAlarms[0].[AlarmName,StateValue,Metrics[2].Expression]' --output text
try-5xx-rate	INSUFFICIENT_DATA	100 * FILL(e, 0) / r

Alarm on symptoms users feel - error rate, p99 latency, the queue's oldest message - rather than on causes such as CPU. A CPU alarm at 80% pages you for a busy but healthy box and stays silent for a deadlock at 2% CPU.

Making it page: SNS

An action is an ARN. For notifications it is an SNS topic, and the topic delivers to its subscriptions: email, SMS, an HTTPS endpoint (PagerDuty, Opsgenie), SQS, Lambda. The shop's on-call topic already exists:

$ aws sns list-subscriptions-by-topic --topic-arn arn:aws:sns:eu-central-1:111122223333:oncall-pages --query 'Subscriptions[].[Protocol,Endpoint,SubscriptionArn]' --output text
email	[email protected]	arn:aws:sns:eu-central-1:111122223333:oncall-pages:28ab2bd4-7eb1-4fa9-abaf-66102563eb6c
$ aws sns create-topic --name try-pages
{
    "TopicArn": "arn:aws:sns:eu-central-1:111122223333:try-pages"
}
$ aws sns subscribe --topic-arn arn:aws:sns:eu-central-1:111122223333:try-pages --protocol email --notification-endpoint [email protected]
{
    "SubscriptionArn": "pending confirmation"
}
$ aws sns list-subscriptions-by-topic --topic-arn arn:aws:sns:eu-central-1:111122223333:try-pages --query 'Subscriptions[].SubscriptionArn' --output text
PendingConfirmation

PendingConfirmation: SNS sent a confirmation email, and until someone clicks its link the subscription receives nothing - the classic "we set up the alarm and the topic, the page never came". The lab's inbox (simulator) is ~/oncall-lab/labs/aws/pager.log; the confirmation is in it, with the token the link carries:

$ tail -4 ~/oncall-lab/labs/aws/pager.log
Confirm subscription

(simulator) the link carries this token - confirm from the CLI with:
  aws sns confirm-subscription --topic-arn arn:aws:sns:eu-central-1:111122223333:try-pages --token 439fd53316671291aa2ada1e7a37495a9c02361be6f3563cf18e747f881a703b
$ aws sns confirm-subscription --topic-arn arn:aws:sns:eu-central-1:111122223333:try-pages --token $(grep -o -- '--token [0-9a-f]*' ~/oncall-lab/labs/aws/pager.log | tail -1 | cut -d' ' -f2)
{
    "SubscriptionArn": "arn:aws:sns:eu-central-1:111122223333:try-pages:62a2486b-a695-4c19-b269-99376e74a85a"
}

Now wire the alarm to it and test the page without breaking anything: set-alarm-state puts an alarm into a state by hand, the action runs, and the next evaluation (within a minute) puts the real state back:

$ aws cloudwatch put-metric-alarm --alarm-name try-5xx --namespace AWS/ApplicationELB --metric-name HTTPCode_Target_5XX_Count --dimensions Name=LoadBalancer,Value=app/shop-alb/50dc6c495c0c9188 --statistic Sum --period 60 --evaluation-periods 3 --datapoints-to-alarm 2 --threshold 5 --comparison-operator GreaterThanThreshold --treat-missing-data notBreaching --alarm-actions arn:aws:sns:eu-central-1:111122223333:try-pages --ok-actions arn:aws:sns:eu-central-1:111122223333:try-pages
$ aws cloudwatch set-alarm-state --alarm-name try-5xx --state-value ALARM --state-reason "testing the page (learner)"
$ grep '^Subject' ~/oncall-lab/labs/aws/pager.log | tail -1
Subject: ALARM: "try-5xx" in EU (Frankfurt)
$ sleep 60
$ aws cloudwatch describe-alarm-history --alarm-name try-5xx --history-item-type StateUpdate --max-items 2 --query 'AlarmHistoryItems[].HistorySummary' --output text
Alarm updated from ALARM to OK	Alarm updated from OK to ALARM

Three more things you will meet:

In an interview: "An alarm never fired during an outage. What do you check?" - "Its history and StateReasonData: was it in INSUFFICIENT_DATA because missing data was left as 'missing' on a sparse metric, were the dimensions exactly the metric's, was M-out-of-N too strict for the period, and did the action work - the SNS topic exists and the subscription is confirmed. Then test with set-alarm-state."

You can now: create an alarm with the right period, M out of N and missing-data treatment, read its state, reason and history, alarm on a rate with metric math, wire it to an SNS topic with a confirmed subscription, and test the page with set-alarm-state.

Why it helps

An alarm that never fires is worse than no alarm: everyone believes they are covered. The two settings behind most silent alarms - what a missing datapoint means, and whether the notification path really works - are what this lesson is about.

Good alarms watch what users feel (error rate, latency), wait for a pattern rather than one bad minute, and are tested end to end. That is the difference between being paged for real problems and learning to ignore the pager.

Commands in this lesson

aws sleep cat tail grep

FAQ

Why is my new alarm stuck in INSUFFICIENT_DATA?

It has no data to decide on. Often the metric is sparse - an error count that is only published when errors happen - and missing data is treated as "missing". Set --treat-missing-data notBreaching for such metrics, check the dimensions match the metric exactly, and remember a new alarm starts in INSUFFICIENT_DATA.

What does "3 out of 5" mean in an alarm?

The alarm looks at the last five periods (--evaluation-periods 5) and goes to ALARM when at least three of them breach the threshold (--datapoints-to-alarm 3). It ignores one or two noisy minutes but still fires within a few minutes of a real problem.

Does an alarm keep notifying while it is in ALARM?

No. Actions run when the state changes, once. If the page is missed, the alarm will not remind anyone; the paging tool handles escalation. That is also why an OK action matters: it tells the on-call engineer that the problem cleared.

How do I test that an alarm really pages?

Use aws cloudwatch set-alarm-state to put it into ALARM by hand. The real action runs - the topic, the subscriptions, the email or pager - and the next evaluation, within a minute, puts the true state back. Do it for every new alarm and for every new person on the rotation.

Why did my email subscription never receive anything?

An email subscription to an SNS topic starts as PendingConfirmation, and SNS delivers nothing to it until someone clicks the confirmation link. list-subscriptions-by-topic shows the state. A topic that does not exist at all makes the alarm's action fail silently, visible only in the alarm history.

In an interview Mid

An alarm never fired during an outage. How do you find out why?

Read its history and StateReasonData with describe-alarm-history: was it sitting in INSUFFICIENT_DATA because missing data was left as "missing" on a sparse metric, were the dimensions exactly the metric's, was M out of N too strict for the period? Then check the action path: the SNS topic exists and the subscription is confirmed, not PendingConfirmation. Finally test the whole path with set-alarm-state.

Also asked: When would you treat missing data as breaching? · Why alarm on an error rate instead of an error count? · How do you stop an alarm from paging during a maintenance window?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.