OnCallReady

Lesson 31.8 · AWS III: CloudWatch, CloudTrail, Cost & Incidents · 18 min read

CloudWatch Logs: groups, retention, filter patterns and tail

In plain words

Every program on AWS writes a diary of what it did. CloudWatch Logs keeps those diaries in folders called log groups, one per service, with a separate notebook (a log stream) for each machine or container.

You can read the newest pages live, search all the notebooks of a folder at once with a search pattern, and ask CloudWatch to count every line that matches a pattern - so that a counter, and an alarm on it, goes up whenever the diary says something bad.

CloudWatch Logs: groups, retention, filter patterns and tail

Metrics tell you that checkout fails; logs tell you why. On AWS, applications, Lambda functions, VPC flow logs, CloudTrail and the CloudWatch agent all ship their lines to CloudWatch Logs. This lesson is the everyday half of it: where the logs are, how long they stay, how to follow them live and how to search them with filter patterns - and how to turn a log line into a metric you can alarm on.

Need to know: a log group holds the logs of one thing (/oncall-lab/shop-api), split into log streams (one per instance, container or pod). A new group keeps events forever until you set a retention (put-retention-policy). aws logs tail GROUP --since 10m --follow is tail -f for a whole group. filter-log-events searches with a filter pattern and takes times in milliseconds. A metric filter counts matching events into a CloudWatch metric - only events that arrive after it exists.

Groups and streams

$ aws logs describe-log-groups --query 'logGroups[].[logGroupName,retentionInDays,storedBytes]' --output table
-------------------------------------------------------------------
|                        DescribeLogGroups                        |
+--------------------------------------------+-------+------------+
|  /aws/vpc/flowlogs/shop-vpc                |  7    |  997920    |
|  /oncall-lab/nginx-access                  |  14   |  15422400  |
|  /oncall-lab/payments                      |  30   |  10108800  |
|  /oncall-lab/shop-api                      |  30   |  40435200  |
|  /oncall-lab/try-debug                     |  None |  31104001  |
|  aws-cloudtrail-logs-111122223333-5f1e2c3d |  400  |  32400004  |
+--------------------------------------------+-------+------------+
$ aws logs describe-log-streams --log-group-name /oncall-lab/shop-api --order-by LastEventTime --descending --query 'logStreams[].[logStreamName,lastEventTimestamp]' --output text
i-0a1b2c3d4e5f60718	1790107203436
i-0b2c3d4e5f6071829	1790107196362

One row has no retention: None in the table is "never expire", the default for every group a service or a person creates. /oncall-lab/try-debug has been collecting debug lines for over a year. Logs cost twice: ingestion (about $0.50 per GB in us-east-1, $0.63 in Frankfurt) when they arrive, and storage (about $0.03 per GB-month, compressed) for as long as you keep them. Ingestion dominates at first; storage without retention grows every month forever. Setting a retention is the cheapest cost fix in AWS:

$ aws logs put-retention-policy --log-group-name /oncall-lab/try-debug --retention-in-days 10
aws: [ERROR]: An error occurred (InvalidParameterException) when calling the PutRetentionPolicy operation: 1 validation error detected: Value '10' at 'retentionInDays' failed to satisfy constraint: Member must satisfy enum value set: [1, 3, 5, 7, 14, 30, 60, 90, 120, 150, 180, 365, 400, 545, 731, 1096, 1827, 2192, 2557, 2922, 3288, 3653]
$ aws logs put-retention-policy --log-group-name /oncall-lab/try-debug --retention-in-days 14
$ aws logs describe-log-groups --log-group-name-prefix /oncall-lab/try --query 'logGroups[].[logGroupName,retentionInDays]' --output text
/oncall-lab/try-debug	14

Only fixed values are allowed (1, 3, 5, 7, 14, 30, 60, 90 ... 3653 days). Logs that must be kept for years (audit, compliance) belong in S3 with a lifecycle rule - export them or let CloudTrail deliver there - not in a log group at CloudWatch prices.

Following a group: aws logs tail

aws logs tail is the command you will type most. It reads all streams of a group, interleaved by time, from --since (default 10 minutes) and with --follow keeps polling for new events until Ctrl+C:

$ aws logs tail /oncall-lab/shop-api --since 1m --format short | head -3
2026-09-22T19:59:05 {"timestamp":"2026-09-22T19:59:05.895Z","level":"INFO","service":"shop-api","msg":"request completed","method":"GET","path":"/api/cart","status":200,"latency_ms":53,"trace_id":"3f72ab04c82870b598a38fa188279475"}
2026-09-22T19:59:05 {"timestamp":"2026-09-22T19:59:05.900Z","level":"INFO","service":"shop-api","msg":"request completed","method":"GET","path":"/api/cart","status":200,"latency_ms":64,"trace_id":"53147bb66983584e504ea0b42e6a1890"}
2026-09-22T19:59:08 {"timestamp":"2026-09-22T19:59:08.342Z","level":"INFO","service":"shop-api","msg":"request completed","method":"POST","path":"/api/checkout","status":200,"latency_ms":195,"trace_id":"ab48665dc01c9d6a013976a388cfb359"}
$ aws logs tail /oncall-lab/shop-api --since 30m --filter-pattern '{ $.status >= 500 }' --format short | head -4
2026-09-22T19:40:18 {"timestamp":"2026-09-22T19:40:18.285Z","level":"ERROR","service":"shop-api","msg":"upstream timeout calling payments","method":"POST","path":"/api/checkout","status":504,"latency_ms":3785,"trace_id":"8543f4a2cd37aa7a464a47ef7f10f73c","error":"ReadTimeout"}
2026-09-22T19:40:20 {"timestamp":"2026-09-22T19:40:20.519Z","level":"ERROR","service":"shop-api","msg":"upstream timeout calling payments","method":"POST","path":"/api/checkout","status":504,"latency_ms":4535,"trace_id":"4450f86c63f425f3496c574aa0557303","error":"ReadTimeout"}
2026-09-22T19:41:24 {"timestamp":"2026-09-22T19:41:24.951Z","level":"ERROR","service":"shop-api","msg":"upstream timeout calling payments","method":"POST","path":"/api/checkout","status":504,"latency_ms":5357,"trace_id":"816b7aa371c1eaffcf38acd0b9b40997","error":"ReadTimeout"}
2026-09-22T19:42:16 {"timestamp":"2026-09-22T19:42:16.767Z","level":"ERROR","service":"shop-api","msg":"upstream timeout calling payments","method":"POST","path":"/api/checkout","status":504,"latency_ms":5936,"trace_id":"6d5412979b745c62449541513d11c6a8","error":"ReadTimeout"}

--format detailed (the default) adds the stream name - which instance or pod said it. --format json pretty-prints JSON messages. On a real incident: aws logs tail /oncall-lab/shop-api --follow --filter-pattern ERROR in one terminal while you work in another.

Filter patterns

The same pattern language is used by filter-log-events, aws logs tail --filter-pattern, metric filters and subscription filters. It is not a regular expression by default:

patternmatches
ERRORevents containing the term ERROR (case-sensitive)
ERROR timeoutboth terms, anywhere
"upstream timeout"the exact phrase
?ERROR ?WARNeither term
ERROR -healthcheckERROR but not healthcheck
{ $.status >= 500 }JSON events whose status field is 500 or more
{ $.level = "ERROR" && $.path = "/api/checkout" }JSON fields combined (also ||, NOT EXISTS, IS NULL)
[ip, id, user, ts, request, status = 5*, bytes, ...]space-delimited fields: the 6th starts with 5
%timeout after [0-9]+ms%a regular expression (between percent signs)
"" (empty)everything

filter-log-events is the API underneath. Its times are milliseconds since the epoch - the most common reason it "finds nothing":

$ aws logs filter-log-events --log-group-name /oncall-lab/shop-api --start-time $(date -d '-30 min' +%s) --end-time $(date +%s) --filter-pattern '{ $.status = 504 }' --query 'length(events)'
0
$ aws logs filter-log-events --log-group-name /oncall-lab/shop-api --start-time $(date -d '-30 min' +%s000) --end-time $(date +%s000) --filter-pattern '{ $.status = 504 }' --query 'length(events)'
16
$ aws logs filter-log-events --log-group-name /oncall-lab/shop-api --start-time $(date -d '-30 min' +%s000) --filter-pattern '{ $.status = 504 }' --max-items 1 --query 'events[0].[logStreamName,message]' --output text
i-0b2c3d4e5f6071829	{"timestamp":"2026-09-22T19:40:18.285Z","level":"ERROR","service":"shop-api","msg":"upstream timeout calling payments","method":"POST","path":"/api/checkout","status":504,"latency_ms":3785,"trace_id":"8543f4a2cd37aa7a464a47ef7f10f73c","error":"ReadTimeout"}

date +%s gives seconds. Read as milliseconds, the first window is 30 minutes on 21 January 1970: zero events and no error. Append 000 (or use $(($(date +%s) * 1000))). Without --start-time the call searches the whole group, page by page - on a big group that takes minutes.

Plain text logs use the space-delimited form. An nginx access line has the client, two dashes, the time in brackets, the request in quotes, the status, the size ...; brackets and quotes group a field:

$ aws logs filter-log-events --log-group-name /oncall-lab/nginx-access --start-time $(date -d '-30 min' +%s000) --filter-pattern '[ip, id, user, ts, request = "POST /api/checkout*", status = 5*, bytes, ref, agent, rt]' --query 'length(events)'
29
$ aws logs tail /oncall-lab/nginx-access --since 30m --filter-pattern '[ip, id, user, ts, request, status = 5*, ...]' --format short | head -2
2026-09-22T19:40:05 10.40.0.17 - - [22/Sep/2026:19:40:05 +0000] "POST /api/checkout HTTP/1.1" 504 186 "-" "Mozilla/5.0 (X11; Linux aarch64)" 5.000
2026-09-22T19:40:26 10.40.1.140 - - [22/Sep/2026:19:40:26 +0000] "POST /api/checkout HTTP/1.1" 504 169 "-" "Mozilla/5.0 (X11; Linux aarch64)" 5.000

... matches any number of fields. test-metric-filter shows what a pattern matches and which values it extracts, without touching a log group - the way to debug one:

$ aws logs test-metric-filter --filter-pattern '[ip, id, user, ts, request, status = 5*, bytes, ...]' --log-event-messages '10.40.0.17 - - [08/Oct/2026:09:41:12 +0000] "POST /api/checkout HTTP/1.1" 502 157 "-" "curl/8.12" 1.002' '10.40.0.17 - - [08/Oct/2026:09:41:13 +0000] "GET /api/cart HTTP/1.1" 200 812 "-" "curl/8.12" 0.031' --query 'matches[].[eventNumber,extractedValues."$status",extractedValues."$request"]' --output text
1	502	POST /api/checkout HTTP/1.1

From logs to metrics: metric filters

A metric filter runs a pattern on every new event of a group and publishes a CloudWatch metric: 1 per match (or a number taken from the event: $.latency_ms). Then you can alarm on "checkout 5xx in the application's own logs" even when no load balancer metric has it:

$ aws logs put-metric-filter --log-group-name /oncall-lab/shop-api --filter-name try-checkout-5xx --filter-pattern '{ $.path = "/api/checkout" && $.status >= 500 }' --metric-transformations metricName=TryCheckout5xx,metricNamespace=Try/Shop,metricValue=1,defaultValue=0
$ aws logs describe-metric-filters --log-group-name /oncall-lab/shop-api --query 'metricFilters[].[filterName,metricTransformations[0].metricNamespace,metricTransformations[0].metricName]' --output text
try-checkout-5xx	Try/Shop	TryCheckout5xx
$ sleep 120
$ aws cloudwatch get-metric-statistics --namespace Try/Shop --metric-name TryCheckout5xx --start-time $(date -u -d '-30 min' +%FT%TZ) --end-time $(date -u +%FT%TZ) --period 300 --statistics Sum --query 'sort_by(Datapoints, &Timestamp)[].[Timestamp,Sum]' --output text
2026-09-22T19:57:00+00:00	0
2026-09-22T20:02:00+00:00	0

The 504s of twenty minutes ago are not in it: a metric filter only sees events ingested after it was created, never the past. And with defaultValue=0 the metric gets a 0 in minutes when events arrived but none matched - a continuous series an alarm can treat normally; without it the metric is as sparse as the ALB's 5xx count.

Subscription filters are the other consumer: they stream every matching event to Lambda, Kinesis or Firehose in near real time (to ship logs to OpenSearch, Splunk or a SIEM). A group can have two.

In an interview: "How would you get alerted on a specific error message in an application's logs on AWS?" - "A metric filter on its log group with a filter pattern for the message (a JSON selector if the logs are structured), publishing a count with defaultValue 0, and a CloudWatch alarm on that metric with an SNS action. It only counts new events, so test the pattern first with test-metric-filter."

You can now: find a service's log group and its active streams, set retention (and explain what logs cost), follow a group with aws logs tail, write filter patterns for text, JSON and space-delimited logs, avoid the milliseconds trap of filter-log-events, and turn log events into an alarmable metric.

Why it helps

When a metric says checkout fails, the logs say why. Knowing where each service's logs are, how to follow them live and how to search them with filter patterns is the everyday half of every incident.

Logs are also one of the classic cost surprises on AWS: every new log group keeps data forever unless you set a retention, and a debug log level in production can cost more than the servers. A retention sweep is one of the cheapest fixes there is.

Commands in this lesson

aws sleep

FAQ

How long does CloudWatch Logs keep my logs?

Forever by default: a new log group has no retention and shows "Never expire". Set one with aws logs put-retention-policy (1, 3, 5, 7, 14, 30 ... 3653 days). Logs that must be kept for years belong in S3 with a lifecycle rule, which is much cheaper.

Why does filter-log-events find nothing although the events are there?

Its --start-time and --end-time are milliseconds since the epoch, not seconds. date +%s gives seconds, which read as milliseconds point to January 1970. Append 000 to the date output. Logs Insights' start-query, on the other hand, takes seconds.

Is a filter pattern a regular expression?

Not by default. Terms are matched as text, quoted phrases exactly, ?term means OR and -term excludes; JSON logs use selectors like { $.status >= 500 }, plain text uses space-delimited fields in brackets. A regular expression goes between percent signs. test-metric-filter shows what a pattern matches.

Why does my new metric filter show nothing for the errors of last night?

A metric filter only counts events that arrive after it was created; it never processes the past. For history, use Logs Insights or filter-log-events. Add defaultValue=0 so the metric publishes zeros when no event matches and the series stays continuous.

How do I follow a service's logs live?

aws logs tail GROUP --follow, optionally with --since 10m, --format short and --filter-pattern. It reads every stream of the group, interleaved by time, and keeps polling until you press Ctrl+C - tail -f for a whole fleet.

In an interview Mid

How would you get alerted on a specific error message in an application's logs on AWS?

Create a metric filter on the application's log group with put-metric-filter: a filter pattern for the message (a JSON selector if the logs are structured), publishing a count with defaultValue 0. Then a CloudWatch alarm on that metric with an SNS action. A metric filter only counts new events, so test the pattern with test-metric-filter first.

Also asked: Why is "Never expire" a problem for CloudWatch Logs? · How do you follow the logs of a whole fleet of instances live? · What is the difference between a metric filter and a subscription filter?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.