CloudWatch Logs: groups, retention, filter patterns and tail
Metrics tell you that checkout fails; logs tell you why. On AWS, applications, Lambda functions, VPC flow logs, CloudTrail and the CloudWatch agent all ship their lines to CloudWatch Logs. This lesson is the everyday half of it: where the logs are, how long they stay, how to follow them live and how to search them with filter patterns - and how to turn a log line into a metric you can alarm on.
Need to know: a log group holds the logs of one thing (/oncall-lab/shop-api), split into log streams (one per instance, container or pod). A new group keeps events forever until you set a retention (put-retention-policy). aws logs tail GROUP --since 10m --follow is tail -f for a whole group. filter-log-events searches with a filter pattern and takes times in milliseconds. A metric filter counts matching events into a CloudWatch metric - only events that arrive after it exists.
Groups and streams
$ aws logs describe-log-groups --query 'logGroups[].[logGroupName,retentionInDays,storedBytes]' --output table
-------------------------------------------------------------------
| DescribeLogGroups |
+--------------------------------------------+-------+------------+
| /aws/vpc/flowlogs/shop-vpc | 7 | 997920 |
| /oncall-lab/nginx-access | 14 | 15422400 |
| /oncall-lab/payments | 30 | 10108800 |
| /oncall-lab/shop-api | 30 | 40435200 |
| /oncall-lab/try-debug | None | 31104001 |
| aws-cloudtrail-logs-111122223333-5f1e2c3d | 400 | 32400004 |
+--------------------------------------------+-------+------------+
$ aws logs describe-log-streams --log-group-name /oncall-lab/shop-api --order-by LastEventTime --descending --query 'logStreams[].[logStreamName,lastEventTimestamp]' --output text
i-0a1b2c3d4e5f60718 1790107203436
i-0b2c3d4e5f6071829 1790107196362
One row has no retention: None in the table is "never expire", the default for every group a service or a person creates. /oncall-lab/try-debug has been collecting debug lines for over a year. Logs cost twice: ingestion (about $0.50 per GB in us-east-1, $0.63 in Frankfurt) when they arrive, and storage (about $0.03 per GB-month, compressed) for as long as you keep them. Ingestion dominates at first; storage without retention grows every month forever. Setting a retention is the cheapest cost fix in AWS:
$ aws logs put-retention-policy --log-group-name /oncall-lab/try-debug --retention-in-days 10
aws: [ERROR]: An error occurred (InvalidParameterException) when calling the PutRetentionPolicy operation: 1 validation error detected: Value '10' at 'retentionInDays' failed to satisfy constraint: Member must satisfy enum value set: [1, 3, 5, 7, 14, 30, 60, 90, 120, 150, 180, 365, 400, 545, 731, 1096, 1827, 2192, 2557, 2922, 3288, 3653]
$ aws logs put-retention-policy --log-group-name /oncall-lab/try-debug --retention-in-days 14
$ aws logs describe-log-groups --log-group-name-prefix /oncall-lab/try --query 'logGroups[].[logGroupName,retentionInDays]' --output text
/oncall-lab/try-debug 14
Only fixed values are allowed (1, 3, 5, 7, 14, 30, 60, 90 ... 3653 days). Logs that must be kept for years (audit, compliance) belong in S3 with a lifecycle rule - export them or let CloudTrail deliver there - not in a log group at CloudWatch prices.
Following a group: aws logs tail
aws logs tail is the command you will type most. It reads all streams of a group, interleaved by time, from --since (default 10 minutes) and with --follow keeps polling for new events until Ctrl+C:
$ aws logs tail /oncall-lab/shop-api --since 1m --format short | head -3
2026-09-22T19:59:05 {"timestamp":"2026-09-22T19:59:05.895Z","level":"INFO","service":"shop-api","msg":"request completed","method":"GET","path":"/api/cart","status":200,"latency_ms":53,"trace_id":"3f72ab04c82870b598a38fa188279475"}
2026-09-22T19:59:05 {"timestamp":"2026-09-22T19:59:05.900Z","level":"INFO","service":"shop-api","msg":"request completed","method":"GET","path":"/api/cart","status":200,"latency_ms":64,"trace_id":"53147bb66983584e504ea0b42e6a1890"}
2026-09-22T19:59:08 {"timestamp":"2026-09-22T19:59:08.342Z","level":"INFO","service":"shop-api","msg":"request completed","method":"POST","path":"/api/checkout","status":200,"latency_ms":195,"trace_id":"ab48665dc01c9d6a013976a388cfb359"}
$ aws logs tail /oncall-lab/shop-api --since 30m --filter-pattern '{ $.status >= 500 }' --format short | head -4
2026-09-22T19:40:18 {"timestamp":"2026-09-22T19:40:18.285Z","level":"ERROR","service":"shop-api","msg":"upstream timeout calling payments","method":"POST","path":"/api/checkout","status":504,"latency_ms":3785,"trace_id":"8543f4a2cd37aa7a464a47ef7f10f73c","error":"ReadTimeout"}
2026-09-22T19:40:20 {"timestamp":"2026-09-22T19:40:20.519Z","level":"ERROR","service":"shop-api","msg":"upstream timeout calling payments","method":"POST","path":"/api/checkout","status":504,"latency_ms":4535,"trace_id":"4450f86c63f425f3496c574aa0557303","error":"ReadTimeout"}
2026-09-22T19:41:24 {"timestamp":"2026-09-22T19:41:24.951Z","level":"ERROR","service":"shop-api","msg":"upstream timeout calling payments","method":"POST","path":"/api/checkout","status":504,"latency_ms":5357,"trace_id":"816b7aa371c1eaffcf38acd0b9b40997","error":"ReadTimeout"}
2026-09-22T19:42:16 {"timestamp":"2026-09-22T19:42:16.767Z","level":"ERROR","service":"shop-api","msg":"upstream timeout calling payments","method":"POST","path":"/api/checkout","status":504,"latency_ms":5936,"trace_id":"6d5412979b745c62449541513d11c6a8","error":"ReadTimeout"}
--format detailed (the default) adds the stream name - which instance or pod said it. --format json pretty-prints JSON messages. On a real incident: aws logs tail /oncall-lab/shop-api --follow --filter-pattern ERROR in one terminal while you work in another.
Filter patterns
The same pattern language is used by filter-log-events, aws logs tail --filter-pattern, metric filters and subscription filters. It is not a regular expression by default:
| pattern | matches |
|---|---|
ERROR | events containing the term ERROR (case-sensitive) |
ERROR timeout | both terms, anywhere |
"upstream timeout" | the exact phrase |
?ERROR ?WARN | either term |
ERROR -healthcheck | ERROR but not healthcheck |
{ $.status >= 500 } | JSON events whose status field is 500 or more |
{ $.level = "ERROR" && $.path = "/api/checkout" } | JSON fields combined (also ||, NOT EXISTS, IS NULL) |
[ip, id, user, ts, request, status = 5*, bytes, ...] | space-delimited fields: the 6th starts with 5 |
%timeout after [0-9]+ms% | a regular expression (between percent signs) |
"" (empty) | everything |
filter-log-events is the API underneath. Its times are milliseconds since the epoch - the most common reason it "finds nothing":
$ aws logs filter-log-events --log-group-name /oncall-lab/shop-api --start-time $(date -d '-30 min' +%s) --end-time $(date +%s) --filter-pattern '{ $.status = 504 }' --query 'length(events)'
0
$ aws logs filter-log-events --log-group-name /oncall-lab/shop-api --start-time $(date -d '-30 min' +%s000) --end-time $(date +%s000) --filter-pattern '{ $.status = 504 }' --query 'length(events)'
16
$ aws logs filter-log-events --log-group-name /oncall-lab/shop-api --start-time $(date -d '-30 min' +%s000) --filter-pattern '{ $.status = 504 }' --max-items 1 --query 'events[0].[logStreamName,message]' --output text
i-0b2c3d4e5f6071829 {"timestamp":"2026-09-22T19:40:18.285Z","level":"ERROR","service":"shop-api","msg":"upstream timeout calling payments","method":"POST","path":"/api/checkout","status":504,"latency_ms":3785,"trace_id":"8543f4a2cd37aa7a464a47ef7f10f73c","error":"ReadTimeout"}
date +%s gives seconds. Read as milliseconds, the first window is 30 minutes on 21 January 1970: zero events and no error. Append 000 (or use $(($(date +%s) * 1000))). Without --start-time the call searches the whole group, page by page - on a big group that takes minutes.
Plain text logs use the space-delimited form. An nginx access line has the client, two dashes, the time in brackets, the request in quotes, the status, the size ...; brackets and quotes group a field:
$ aws logs filter-log-events --log-group-name /oncall-lab/nginx-access --start-time $(date -d '-30 min' +%s000) --filter-pattern '[ip, id, user, ts, request = "POST /api/checkout*", status = 5*, bytes, ref, agent, rt]' --query 'length(events)'
29
$ aws logs tail /oncall-lab/nginx-access --since 30m --filter-pattern '[ip, id, user, ts, request, status = 5*, ...]' --format short | head -2
2026-09-22T19:40:05 10.40.0.17 - - [22/Sep/2026:19:40:05 +0000] "POST /api/checkout HTTP/1.1" 504 186 "-" "Mozilla/5.0 (X11; Linux aarch64)" 5.000
2026-09-22T19:40:26 10.40.1.140 - - [22/Sep/2026:19:40:26 +0000] "POST /api/checkout HTTP/1.1" 504 169 "-" "Mozilla/5.0 (X11; Linux aarch64)" 5.000
... matches any number of fields. test-metric-filter shows what a pattern matches and which values it extracts, without touching a log group - the way to debug one:
$ aws logs test-metric-filter --filter-pattern '[ip, id, user, ts, request, status = 5*, bytes, ...]' --log-event-messages '10.40.0.17 - - [08/Oct/2026:09:41:12 +0000] "POST /api/checkout HTTP/1.1" 502 157 "-" "curl/8.12" 1.002' '10.40.0.17 - - [08/Oct/2026:09:41:13 +0000] "GET /api/cart HTTP/1.1" 200 812 "-" "curl/8.12" 0.031' --query 'matches[].[eventNumber,extractedValues."$status",extractedValues."$request"]' --output text
1 502 POST /api/checkout HTTP/1.1
From logs to metrics: metric filters
A metric filter runs a pattern on every new event of a group and publishes a CloudWatch metric: 1 per match (or a number taken from the event: $.latency_ms). Then you can alarm on "checkout 5xx in the application's own logs" even when no load balancer metric has it:
$ aws logs put-metric-filter --log-group-name /oncall-lab/shop-api --filter-name try-checkout-5xx --filter-pattern '{ $.path = "/api/checkout" && $.status >= 500 }' --metric-transformations metricName=TryCheckout5xx,metricNamespace=Try/Shop,metricValue=1,defaultValue=0
$ aws logs describe-metric-filters --log-group-name /oncall-lab/shop-api --query 'metricFilters[].[filterName,metricTransformations[0].metricNamespace,metricTransformations[0].metricName]' --output text
try-checkout-5xx Try/Shop TryCheckout5xx
$ sleep 120
$ aws cloudwatch get-metric-statistics --namespace Try/Shop --metric-name TryCheckout5xx --start-time $(date -u -d '-30 min' +%FT%TZ) --end-time $(date -u +%FT%TZ) --period 300 --statistics Sum --query 'sort_by(Datapoints, &Timestamp)[].[Timestamp,Sum]' --output text
2026-09-22T19:57:00+00:00 0
2026-09-22T20:02:00+00:00 0
The 504s of twenty minutes ago are not in it: a metric filter only sees events ingested after it was created, never the past. And with defaultValue=0 the metric gets a 0 in minutes when events arrived but none matched - a continuous series an alarm can treat normally; without it the metric is as sparse as the ALB's 5xx count.
Subscription filters are the other consumer: they stream every matching event to Lambda, Kinesis or Firehose in near real time (to ship logs to OpenSearch, Splunk or a SIEM). A group can have two.
In an interview: "How would you get alerted on a specific error message in an application's logs on AWS?" - "A metric filter on its log group with a filter pattern for the message (a JSON selector if the logs are structured), publishing a count with defaultValue 0, and a CloudWatch alarm on that metric with an SNS action. It only counts new events, so test the pattern first with test-metric-filter."
You can now: find a service's log group and its active streams, set retention (and explain what logs cost), follow a group with aws logs tail, write filter patterns for text, JSON and space-delimited logs, avoid the milliseconds trap of filter-log-events, and turn log events into an alarmable metric.