OnCallReady

Lesson 31.11 · AWS III: CloudWatch, CloudTrail, Cost & Incidents · 20 min read

Logs Insights: query your logs like a database

In plain words

Searching logs line by line is like reading every page of every diary to count how often someone wrote "error". Logs Insights is a librarian you can ask questions: "how many errors per message in the last hour?", "when did they start?", "which page is the slowest?".

You write the question in a small query language - keep these lines, count them by this field, sort the counts - and Insights reads all the diaries of a folder at once and hands you a little table.

Logs Insights: query your logs like a database

Filter patterns find lines. During an incident you need answers: which error, on which path, since when, how many per minute, how slow. CloudWatch Logs Insights is a query language over log groups - filter, stats ... by, sort, parse - with the fields of JSON logs discovered for you. If you did the Azure chapters, it is KQL's little sibling; if not, think grep | awk | sort | uniq -c that runs over every stream of a group at once.

Need to know: a query is commands joined by |: fields, filter, stats, sort, limit, parse, display, dedup. Every event has @timestamp, @message, @logStream and @log; JSON fields (status, path, latency_ms) are discovered automatically. From the CLI it is two calls: start-query (times in seconds, the query in --query-string) returns a queryId, and get-query-results returns Running until it is Complete. You pay per GB scanned, so narrow the time window first.

Two calls and a status

$ Q=$(aws logs start-query --log-group-name /oncall-lab/shop-api --start-time $(date -d '-15 min' +%s) --end-time $(date +%s) --query-string 'fields @timestamp, path, status | sort @timestamp desc | limit 2' --query queryId --output text); echo $Q
b74b3cd2-8236-4612-97ab-fe297c896783
$ aws logs get-query-results --query-id $Q --query status
"Running"
$ sleep 2
$ aws logs get-query-results --query-id $Q
{
    "queryLanguage": "CWLI",
    "results": [
        [
            {
                "field": "@timestamp",
                "value": "2026-09-22 20:00:03.436"
            },
            {
                "field": "path",
                "value": "/api/cart"
            },
            {
                "field": "status",
                "value": "200"
            },
            {
                "field": "@ptr",
                "value": "CmZXUtY2VudHJhbC0xfC9vbmNhbGwtbGFiL3Nob3AtYXBpfGktMGExYjJjM2Q0ZTVmNjA3MTh8MTc5MDEwNzIwMzQzNnwyODY5NzI4NTI3EAA"
            }
        ],
        [
            {
                "field": "@timestamp",
                "value": "2026-09-22 19:59:56.362"
            },
            {
                "field": "path",
                "value": "/api/products"
            },
            {
                "field": "status",
                "value": "200"
            },
            {
                "field": "@ptr",
                "value": "CmZXUtY2VudHJhbC0xfC9vbmNhbGwtbGFiL3Nob3AtYXBpfGktMGIyYzNkNGU1ZjYwNzE4Mjl8MTc5MDEwNzE5NjM2Mnw0MjEzNDQyMTMxEAA"
            }
        ]
    ],
    "statistics": {
        "recordsMatched": 366,
        "recordsScanned": 366,
        "estimatedRecordsSkipped": 0,
        "bytesScanned": 88187,
        "estimatedBytesSkipped": 0,
        "logGroupsScanned": 1
    },
    "status": "Complete"
}

Read the shape before you script it:

Two flags are easy to confuse: --query-string is the Logs Insights query, --query is the CLI's own JMESPath filter on the response. The JMESPath results[*][*].value with --output text turns the rows into tab-separated lines - wrapped in a shell function, that is the everyday tool:

$ lq() { local q; q=$(aws logs start-query --log-group-name "$1" --start-time $(date -d "-${3:-60} min" +%s) --end-time $(date +%s) --query-string "$2" --query queryId --output text) && sleep 2 && aws logs get-query-results --query-id "$q" --query 'results[*][*].value' --output text; }
$ lq /oncall-lab/shop-api 'stats count(*) as requests by path | sort requests desc'
/api/products	553
/api/cart	373
/api/cart/items	217
/api/orders	143
/api/health	125
/api/checkout	100

The top errors, the timeline, the slow paths

Three queries answer most incidents. What fails:

$ lq /oncall-lab/shop-api 'filter level = "ERROR" | stats count(*) as errors by msg, path | sort errors desc | limit 10'
upstream connect error calling payments	/api/checkout	13

Since when - bin() groups by time, like date_trunc:

$ lq /oncall-lab/shop-api 'filter status >= 500 | stats count(*) as errors by bin(5m)' 60
2026-09-22 19:25:00.000	1
2026-09-22 19:30:00.000	2
2026-09-22 19:35:00.000	2
2026-09-22 19:40:00.000	3
2026-09-22 19:45:00.000	1
2026-09-22 19:50:00.000	2
2026-09-22 19:55:00.000	2
2026-09-22 20:00:00.000	1

The errors start at a clean five-minute boundary and continue to now: that start time is what you put next to the deploy log and the CloudTrail events (the next lesson). What is slow - percentiles per path, the query behind every latency incident:

$ lq /oncall-lab/shop-api 'stats count(*) as n, avg(latency_ms) as avg_ms, pct(latency_ms, 99) as p99_ms by path | sort p99_ms desc'
/api/orders	144	314.625	1156
/api/checkout	100	211.61	334
/api/cart/items	218	59.7889908256881	98
/api/cart	372	44.2983870967742	73
/api/products	553	30.9403254972875	50
/api/health	125	3.176	5

/api/orders has a p99 far above its average: something made a fraction of its requests slow - a lock, a cold cache, one bad instance. Add @logStream to the by and you see whether it is one instance.

The language

commandwhat it doesexample
fieldschoose and compute fieldsfields @timestamp, path, latency_ms / 1000 as s
filterkeep matching eventsfilter status >= 500 and path like "/api/"
statsaggregate, optionally by fields or bin()stats avg(latency_ms) by bin(1m)
sortorder (asc is the default)sort @timestamp desc
limitthe first N rowslimit 20
parseextract fields from textparse @message "user=* " as user
displaywhich fields to showdisplay @timestamp, msg
dedupone row per valuededup path

Operators: =, !=, <, >=, and, or, not, in ["a", "b"], like "text" (substring), like /regex/ and =~ /regex/. Aggregations: count(*), count(field), count_distinct, sum, avg, min, max, pct(field, 99), stddev, earliest, latest. Functions: strlen, tolower, concat, replace, abs, floor, datefloor, ispresent, coalesce.

A syntax error comes back from start-query itself - before anything runs - in the parser's own words:

$ aws logs start-query --log-group-name /oncall-lab/shop-api --start-time $(date -d '-15 min' +%s) --end-time $(date +%s) --query-string 'filter level = "ERROR" | count(*) by msg'
aws: [ERROR]: An error occurred (MalformedQueryException) when calling the StartQuery operation: mismatched input 'count' expecting {K_PARSE, K_SEARCH, K_FIELDS, K_DISPLAY, K_FILTER, K_STATS, K_SORT, K_ORDER, K_HEAD, K_LIMIT, K_TAIL, K_DEDUP, K_UNMASK, K_PATTERN}

Additional error details:
queryCompileError: <complex value>

"mismatched input 'count' expecting {K_PARSE, ...}": count(*) is an aggregation, it needs stats in front of it. The list in braces is what the parser would have accepted there.

Logs that are not JSON: parse

Plain text needs parse to become fields. A glob pattern (* = anything) with one name per *:

$ lq /oncall-lab/nginx-access 'parse @message "* - - [*] \"* * *\" * * \"*\" \"*\" *" as client, ts, method, url, proto, status, bytes, ref, agent, rt | filter status >= 500 | stats count(*) as n, avg(rt) as avg_s by method, url, status | sort n desc' 60
POST	/api/checkout	502	49	0.002

Or a regular expression with named groups, which is stricter about what it accepts:

$ lq /oncall-lab/nginx-access 'parse @message /"(?<method>[A-Z]+) (?<url>[^ ?]+)[^"]*" (?<status>\d{3})/ | stats count(*) as n by status | sort n desc' 30
200	889
502	37

Parsed values are strings that look like numbers; comparisons such as status >= 500 and avg(rt) treat them as numbers.

Several groups, and what you pay

--log-group-names takes up to 50 groups; @log says which one an event came from:

$ Q=$(aws logs start-query --log-group-names /oncall-lab/shop-api /oncall-lab/payments --start-time $(date -d '-30 min' +%s) --end-time $(date +%s) --query-string 'filter level = "ERROR" | stats count(*) as errors by @log' --query queryId --output text); sleep 2
$ aws logs get-query-results --query-id $Q --query '[results[*][*].value, statistics]'
[
    [
        [
            "111122223333:/oncall-lab/shop-api",
            "13"
        ]
    ],
    {
        "recordsMatched": 13,
        "recordsScanned": 925,
        "estimatedRecordsSkipped": 0,
        "bytesScanned": 222819,
        "estimatedBytesSkipped": 0,
        "logGroupsScanned": 2
    }
]

Logs Insights costs about $0.005 per GB scanned (a little more in Frankfurt), and a query scans everything in the time window of every group you name - the filter does not reduce the scan. A 30-day window over a busy group can cost dollars per query and take minutes. Start with the last hour, find the start time, then widen.

In an interview: "Users report errors since this morning. How do you find out what is failing from the logs?" - "A Logs Insights query over the service's log group for the last hours: filter on errors (status >= 500 or level = ERROR), stats count() by message and path to see what fails, then stats count() by bin(5m) to see when it started. That start time is what I correlate with deploys and CloudTrail."

You can now: run a Logs Insights query from the CLI and poll it to Complete, read its result shape and statistics, write the three incident queries (top errors, timeline with bin(), percentiles per path), parse text logs with globs and regexes, query several groups, and keep the scan (and the bill) small.

Why it helps

During an incident you need answers in minutes, not lines to scroll through. Three Logs Insights queries - the top errors, a timeline with bin(), and percentiles per path - answer most of "what is failing, since when and where".

Insights is also what makes CloudTrail and VPC flow logs useful at scale, because it discovers their fields automatically. And because you pay per gigabyte scanned, knowing how to keep the window small matters on a real account.

Commands in this lesson

aws sleep

FAQ

How do I run a Logs Insights query from the CLI?

aws logs start-query with the log group, --start-time and --end-time in epoch seconds and the query in --query-string; it returns a queryId. Then aws logs get-query-results --query-id until the status is Complete. Do not confuse --query-string (the Insights query) with --query (the CLI's JMESPath filter).

Do I have to parse JSON logs first?

No. Insights discovers the fields of JSON events automatically, nested ones with dots (http.status, userIdentity.arn), and also the fields of VPC flow logs and CloudTrail records. Plain text needs the parse command, with a glob pattern or a regular expression with named groups.

Why does my stats query fail with "mismatched input"?

The parser expected a command where you wrote something else - typically an aggregation such as count(*) without stats in front of it. The error lists the commands it would have accepted at that position.

What does a Logs Insights query cost?

About half a cent per GB scanned, and a query scans everything in the time window of every log group you name - the filter does not reduce the scan. Start with a narrow window, find the start time, then widen it.

Is Logs Insights the same as KQL?

They are close cousins: filter is KQL's where, stats ... by is summarize ... by, bin(5m) is bin(TimeGenerated, 5m). If you know one, the other takes an afternoon. The results come back as rows of field/value pairs, with every value a string.

In an interview Mid

Users report errors since this morning. How do you find out what is failing from the logs?

A Logs Insights query over the service's log group: filter status >= 500 (or level = ERROR) | stats count(*) by msg, path | sort desc to see what fails, then stats count(*) by bin(5m) to see when it started, and pct(latency_ms, 99) by path for what is slow. From the CLI that is start-query and get-query-results until Complete. The start time is what I correlate with deploys and CloudTrail.

Also asked: How would you find which of two services started failing first? · How do you extract fields from plain-text logs in Logs Insights? · How do you keep Logs Insights queries cheap?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.