Logs Insights: query your logs like a database
Filter patterns find lines. During an incident you need answers: which error, on which path, since when, how many per minute, how slow. CloudWatch Logs Insights is a query language over log groups - filter, stats ... by, sort, parse - with the fields of JSON logs discovered for you. If you did the Azure chapters, it is KQL's little sibling; if not, think grep | awk | sort | uniq -c that runs over every stream of a group at once.
Need to know: a query is commands joined by |: fields, filter, stats, sort, limit, parse, display, dedup. Every event has @timestamp, @message, @logStream and @log; JSON fields (status, path, latency_ms) are discovered automatically. From the CLI it is two calls: start-query (times in seconds, the query in --query-string) returns a queryId, and get-query-results returns Running until it is Complete. You pay per GB scanned, so narrow the time window first.
Two calls and a status
$ Q=$(aws logs start-query --log-group-name /oncall-lab/shop-api --start-time $(date -d '-15 min' +%s) --end-time $(date +%s) --query-string 'fields @timestamp, path, status | sort @timestamp desc | limit 2' --query queryId --output text); echo $Q
b74b3cd2-8236-4612-97ab-fe297c896783
$ aws logs get-query-results --query-id $Q --query status
"Running"
$ sleep 2
$ aws logs get-query-results --query-id $Q
{
"queryLanguage": "CWLI",
"results": [
[
{
"field": "@timestamp",
"value": "2026-09-22 20:00:03.436"
},
{
"field": "path",
"value": "/api/cart"
},
{
"field": "status",
"value": "200"
},
{
"field": "@ptr",
"value": "CmZXUtY2VudHJhbC0xfC9vbmNhbGwtbGFiL3Nob3AtYXBpfGktMGExYjJjM2Q0ZTVmNjA3MTh8MTc5MDEwNzIwMzQzNnwyODY5NzI4NTI3EAA"
}
],
[
{
"field": "@timestamp",
"value": "2026-09-22 19:59:56.362"
},
{
"field": "path",
"value": "/api/products"
},
{
"field": "status",
"value": "200"
},
{
"field": "@ptr",
"value": "CmZXUtY2VudHJhbC0xfC9vbmNhbGwtbGFiL3Nob3AtYXBpfGktMGIyYzNkNGU1ZjYwNzE4Mjl8MTc5MDEwNzE5NjM2Mnw0MjEzNDQyMTMxEAA"
}
]
],
"statistics": {
"recordsMatched": 366,
"recordsScanned": 366,
"estimatedRecordsSkipped": 0,
"bytesScanned": 88187,
"estimatedBytesSkipped": 0,
"logGroupsScanned": 1
},
"status": "Complete"
}
Read the shape before you script it:
statusgoesScheduled/Running->Complete(orFailed,Cancelled,Timeout). Results read while it runs are partial or empty; poll until Complete;resultsis a list of rows, and each row is a list of{field, value}pairs - every value is a string, numbers included;@ptrpoints at the original event (the console uses it to show the whole record);statisticssays how much was matched and scanned: the bill and the speed both followbytesScanned.
Two flags are easy to confuse: --query-string is the Logs Insights query, --query is the CLI's own JMESPath filter on the response. The JMESPath results[*][*].value with --output text turns the rows into tab-separated lines - wrapped in a shell function, that is the everyday tool:
$ lq() { local q; q=$(aws logs start-query --log-group-name "$1" --start-time $(date -d "-${3:-60} min" +%s) --end-time $(date +%s) --query-string "$2" --query queryId --output text) && sleep 2 && aws logs get-query-results --query-id "$q" --query 'results[*][*].value' --output text; }
$ lq /oncall-lab/shop-api 'stats count(*) as requests by path | sort requests desc'
/api/products 553
/api/cart 373
/api/cart/items 217
/api/orders 143
/api/health 125
/api/checkout 100
The top errors, the timeline, the slow paths
Three queries answer most incidents. What fails:
$ lq /oncall-lab/shop-api 'filter level = "ERROR" | stats count(*) as errors by msg, path | sort errors desc | limit 10'
upstream connect error calling payments /api/checkout 13
Since when - bin() groups by time, like date_trunc:
$ lq /oncall-lab/shop-api 'filter status >= 500 | stats count(*) as errors by bin(5m)' 60
2026-09-22 19:25:00.000 1
2026-09-22 19:30:00.000 2
2026-09-22 19:35:00.000 2
2026-09-22 19:40:00.000 3
2026-09-22 19:45:00.000 1
2026-09-22 19:50:00.000 2
2026-09-22 19:55:00.000 2
2026-09-22 20:00:00.000 1
The errors start at a clean five-minute boundary and continue to now: that start time is what you put next to the deploy log and the CloudTrail events (the next lesson). What is slow - percentiles per path, the query behind every latency incident:
$ lq /oncall-lab/shop-api 'stats count(*) as n, avg(latency_ms) as avg_ms, pct(latency_ms, 99) as p99_ms by path | sort p99_ms desc'
/api/orders 144 314.625 1156
/api/checkout 100 211.61 334
/api/cart/items 218 59.7889908256881 98
/api/cart 372 44.2983870967742 73
/api/products 553 30.9403254972875 50
/api/health 125 3.176 5
/api/orders has a p99 far above its average: something made a fraction of its requests slow - a lock, a cold cache, one bad instance. Add @logStream to the by and you see whether it is one instance.
The language
| command | what it does | example |
|---|---|---|
fields | choose and compute fields | fields @timestamp, path, latency_ms / 1000 as s |
filter | keep matching events | filter status >= 500 and path like "/api/" |
stats | aggregate, optionally by fields or bin() | stats avg(latency_ms) by bin(1m) |
sort | order (asc is the default) | sort @timestamp desc |
limit | the first N rows | limit 20 |
parse | extract fields from text | parse @message "user=* " as user |
display | which fields to show | display @timestamp, msg |
dedup | one row per value | dedup path |
Operators: =, !=, <, >=, and, or, not, in ["a", "b"], like "text" (substring), like /regex/ and =~ /regex/. Aggregations: count(*), count(field), count_distinct, sum, avg, min, max, pct(field, 99), stddev, earliest, latest. Functions: strlen, tolower, concat, replace, abs, floor, datefloor, ispresent, coalesce.
A syntax error comes back from start-query itself - before anything runs - in the parser's own words:
$ aws logs start-query --log-group-name /oncall-lab/shop-api --start-time $(date -d '-15 min' +%s) --end-time $(date +%s) --query-string 'filter level = "ERROR" | count(*) by msg'
aws: [ERROR]: An error occurred (MalformedQueryException) when calling the StartQuery operation: mismatched input 'count' expecting {K_PARSE, K_SEARCH, K_FIELDS, K_DISPLAY, K_FILTER, K_STATS, K_SORT, K_ORDER, K_HEAD, K_LIMIT, K_TAIL, K_DEDUP, K_UNMASK, K_PATTERN}
Additional error details:
queryCompileError: <complex value>
"mismatched input 'count' expecting {K_PARSE, ...}": count(*) is an aggregation, it needs stats in front of it. The list in braces is what the parser would have accepted there.
Logs that are not JSON: parse
Plain text needs parse to become fields. A glob pattern (* = anything) with one name per *:
$ lq /oncall-lab/nginx-access 'parse @message "* - - [*] \"* * *\" * * \"*\" \"*\" *" as client, ts, method, url, proto, status, bytes, ref, agent, rt | filter status >= 500 | stats count(*) as n, avg(rt) as avg_s by method, url, status | sort n desc' 60
POST /api/checkout 502 49 0.002
Or a regular expression with named groups, which is stricter about what it accepts:
$ lq /oncall-lab/nginx-access 'parse @message /"(?<method>[A-Z]+) (?<url>[^ ?]+)[^"]*" (?<status>\d{3})/ | stats count(*) as n by status | sort n desc' 30
200 889
502 37
Parsed values are strings that look like numbers; comparisons such as status >= 500 and avg(rt) treat them as numbers.
Several groups, and what you pay
--log-group-names takes up to 50 groups; @log says which one an event came from:
$ Q=$(aws logs start-query --log-group-names /oncall-lab/shop-api /oncall-lab/payments --start-time $(date -d '-30 min' +%s) --end-time $(date +%s) --query-string 'filter level = "ERROR" | stats count(*) as errors by @log' --query queryId --output text); sleep 2
$ aws logs get-query-results --query-id $Q --query '[results[*][*].value, statistics]'
[
[
[
"111122223333:/oncall-lab/shop-api",
"13"
]
],
{
"recordsMatched": 13,
"recordsScanned": 925,
"estimatedRecordsSkipped": 0,
"bytesScanned": 222819,
"estimatedBytesSkipped": 0,
"logGroupsScanned": 2
}
]
Logs Insights costs about $0.005 per GB scanned (a little more in Frankfurt), and a query scans everything in the time window of every group you name - the filter does not reduce the scan. A 30-day window over a busy group can cost dollars per query and take minutes. Start with the last hour, find the start time, then widen.
In an interview: "Users report errors since this morning. How do you find out what is failing from the logs?" - "A Logs Insights query over the service's log group for the last hours: filter on errors (status >= 500 or level = ERROR), stats count() by message and path to see what fails, then stats count() by bin(5m) to see when it started. That start time is what I correlate with deploys and CloudTrail."
You can now: run a Logs Insights query from the CLI and poll it to Complete, read its result shape and statistics, write the three incident queries (top errors, timeline with bin(), percentiles per path), parse text logs with globs and regexes, query several groups, and keep the scan (and the bill) small.