OnCallReady

Lesson 29.6 · Observability III: Dashboards, Logs, Traces & Incidents · 18 min read

Loki and LogQL

In plain words

Imagine a huge library where books are not indexed by every word inside them, only by shelf labels: "science, year 2026, room 3". Finding the right shelf is instant. Finding a sentence means taking the books off that shelf and reading them quickly. So the trick is always to go to the smallest shelf first, then read.

Loki stores logs like that. A stream is a set of labels, such as {unit="orders.service", host="oncall-lab"}; only labels are indexed, line content is scanned at query time. LogQL starts with a stream selector, then filters lines (|= "ERROR"), then parses them (| json) into temporary labels you can filter on, and can turn logs into numbers with count_over_time. On this box, Grafana Alloy ships the journal to Loki.

Why logs need their own database

The metrics say checkout started failing at 20:01. They cannot say why: that is in the exception the app logged. On this one box journalctl -u orders finds it. On forty servers, or with the container that logged it already gone, there is no box to ssh into. You need every log line from every machine in one place, searchable by the same labels you use in Prometheus.

What you need to know already: journald and journalctl -u (2.30, 2.32); structured JSON logs and trace ids (27.1); labels, series and cardinality (27.1, 27.2); PromQL selectors, matchers (=, !=, =~), rate(), sum by and topk (27.8); regular expressions (7.3); jq (7.11).

The words you need first

Indexed by labels, not by content

Loki indexes the labels of each stream; the text of each line is only compressed into chunks and read when you query. Consequences:

LogQL, part 1: pick streams, then filter lines

Every log query starts with a stream selector in curly braces, the same syntax as a PromQL selector. Then optional line filters, which keep or drop lines by their text:

{unit="orders.service"}                                  stream selector (required)
{unit="orders.service"} |= "ERROR"                       keep lines that contain ERROR
{unit="orders.service"} != "health"                      drop lines that contain health
{unit="orders.service"} |~ "Hikari.*timed out"           keep lines matching a regex
{unit="orders.service"} !~ "DEBUG|TRACE"                 drop lines matching a regex

Read |= as "pipe into: contains", |~ as "contains a regex match"; the ! versions invert. Filters run left to right and are the cheapest thing you can do to content, so put the one that throws away the most lines first.

logcli is Loki's command-line client, the terminal version of Grafana's log view:

$ logcli query '{unit="orders.service"} |= "ERROR"' --limit=2
http://localhost:3100/loki/api/v1/query_range?direction=BACKWARD&end=...&limit=2&query=...
Common labels: {host="oncall-lab", job="systemd-journal", unit="orders.service"}
2026-09-22T20:02:24Z {} {"@timestamp":"2026-09-22T20:02:24.000Z","@version":"1","message":"Request processing failed: ...","level":"ERROR",...}
2026-09-22T20:02:20Z {} {"@timestamp":"2026-09-22T20:02:20.000Z",...}

The output: the first two lines go to stderr - the URL logcli called (Loki's HTTP API) and the common labels every result shares. Then one line per log entry: its timestamp, the labels that are not common ({} = none left over), and the line itself - here a JSON document.

LogQL, part 2: parsers turn text into labels

{unit="orders.service"} | json

A parser reads each line and turns its fields into labels - only for the rest of this one query; nothing is indexed or stored. | json parses JSON lines. Field names are cleaned up to be valid label names: @timestamp becomes _timestamp, http.uri becomes http_uri. | logfmt does the same for logfmt lines, the key=value key=value style Prometheus itself logs in.

After a parser, two more pipeline stages become useful:

{unit="orders.service"} | json | level="ERROR" | line_format "{{.traceId}} {{.message}}"
2026-09-22T20:02:24Z {} 4db9ca9a4eb9cc2d4bb9c7744cb9c907 Request processing failed: org.springframework.jdbc.CannotGetJdbcConnectionException: Failed to obtain JDBC Connection

Piece by piece: select the orders stream, parse every line as JSON, keep the ones whose level field is ERROR, print only the trace id and message. The output is one error per line, easy to read.

Label filters compare numbers too: | json | duration_ms > 1000, | logfmt | status >= 500. A line the parser could not read gets a special label __error__; | __error__="" keeps only the lines that parsed.

LogQL, part 3: metric queries turn logs into numbers

Wrap a log query and a time range in a range function and the result is no longer lines but numbers, exactly like a PromQL vector (27.8):

count_over_time({unit="orders.service"} |= "ERROR" [5m])          number of lines in the last 5m, per stream
rate({unit="orders.service"} |= "ERROR" [5m])                     lines per second
sum by (level) (count_over_time({unit="orders.service"} | json [15m]))
topk(3, sum by (http_uri) (count_over_time({unit="orders.service"} | json | level="ERROR" [1h])))

Take the third one apart, inside out:

The fourth adds topk(3, ...): keep the three biggest, i.e. the three endpoints that logged the most errors in the last hour.

logcli instant-query runs a metric query at one moment (now) and prints JSON:

$ logcli instant-query 'sum by (level) (count_over_time({unit="orders.service"} | json [15m]))'
[
  {
    "metric": { "level": "ERROR" },
    "value": [1790107400, "38"]
  },
  {
    "metric": { "level": "WARN" },
    "value": [1790107400, "11"]
  }
]

One object per result: metric holds its labels, value is [Unix time, "value as a string"] - 38 ERROR and 11 WARN lines in the last 15 minutes.

sum by (level) works on a label that came from | json because the parser created it. This is also how you alert on logs when there is no metric: Loki's ruler evaluates LogQL rules like Prometheus does and sends to the same Alertmanager (28.7). Prefer a real metric when there is one: counting log lines is more expensive and silently breaks when someone rewords a message.

From a log line to a trace

The traceId in each JSON line (27.1) is the bridge to tracing, the next lesson. In Grafana, a "derived field" on the Loki data source turns it into a link that opens the trace. Here you copy it into trace <id> (simulator). The opposite direction - from a trace to the logs of that one request - is a line filter on the id:

{unit="orders.service"} |= "4db9ca9a4eb9cc2d4bb9c7744cb9c907"

That is why the trace id belongs in the line, never in a label: a filter finds it cheaply, while a label would create a stream per request.

journalctl is still there

On one box, journalctl -u orders -o cat | jq -r 'select(.level=="ERROR") | .message' answers the same questions without Loki (-o cat prints only the message, 2.32). What Loki adds is every box and every container in one query, and retention after the container is gone: in Kubernetes the logs of a crashed pod's previous container (kubectl logs --previous) disappear the moment the pod is deleted.

What you can now do

Why it helps

In a cluster, logs of a crashed pod's previous container vanish once the pod is deleted, and on a fleet you cannot journalctl every box. Loki gives you every box and every pod in one query, with retention. During an incident you need {unit="orders.service"} | json | level="ERROR" | line_format "{{.traceId}} {{.message}}" fast, and then the trace id to jump to the trace.

On a platform team you will also run it, and the cardinality rule is the same as Prometheus: a trace id or user id as a label creates a stream per request and takes Loki down. Knowing that Promtail reached end of life in early 2026 and Alloy replaced it matters for any new setup or migration ticket. And alerting on log counts via the Loki ruler is useful when a service has no metric for a failure, though a real metric is better.

Commands in this lesson

logcli

FAQ

Why should I narrow by labels before filtering content?

Loki indexes only stream labels. A stream selector picks which compressed chunks to read, which is cheap. Line filters like |= "ERROR" then scan every line in those chunks over the time range. A broad selector over a wide range means scanning huge amounts of data, which is slow and expensive. Select the smallest set of streams, such as one unit, namespace or app, then filter, putting the most selective filter first.

Why is a trace id not a Loki label?

Every unique combination of label values is a separate stream with its own chunks and index entries. A trace id, request id or user id has a new value for almost every line, so each line would become its own stream, exploding the index and slowing everything down. Keep such values in the line content, where a filter like |= "4db9ca9a..." finds them, or use structured metadata in newer Loki versions (key/value pairs stored next to each line without creating a stream).

What does | json do?

It parses each log line as JSON at query time and turns its fields into labels for the rest of that query only; nothing is indexed. Field names are sanitised: @timestamp becomes _timestamp, nested http.uri becomes http_uri. You can then filter (level="ERROR", duration_ms > 1000), group in metric queries or reformat with line_format. Lines that fail to parse get an __error__ label.

Can I alert on logs?

Yes. Loki's ruler evaluates LogQL metric queries such as sum(rate({unit="orders.service"} |= "ERROR" [5m])) > 1 as alerting rules and sends alerts to Alertmanager, like Prometheus. It is useful when a failure has no metric. Prefer a real metric where possible: log-based alerts cost more to evaluate and break silently when someone rewords a log message.

What replaced Promtail?

Grafana Alloy. Promtail was deprecated in early 2025 and reached end of life in early 2026. Alloy is Grafana's telemetry collector (built on the tracing collector program you meet in 29.11) with its own configuration language; it can read the systemd journal, files and Kubernetes pod logs and ship to Loki, and also handles metrics and traces. On oncall-lab its config is /etc/alloy/config.alloy. Migration tooling exists to convert Promtail configs.

In an interview Mid

What is Loki and how is it different from other log systems?

Loki is a log database that indexes only labels, not the text of the lines. Lines are grouped into streams (all lines with the same label set, e.g. {job="systemd-journal", unit="orders.service", host="oncall-lab"}), compressed into chunks, and only read when a query needs them. A shipper (Grafana Alloy) sends them from each machine.

Consequences:

LogQL looks like PromQL: {unit="orders.service"} |= "ERROR", then a parser (| json, | logfmt) that turns fields into labels for that query, label filters (| level="ERROR"), line_format. Wrap it in count_over_time(...[5m]) or rate and it returns numbers you can sum by - but prefer a real metric when there is one.

Also asked: Write a LogQL query to find the endpoints producing the most errors in the last hour, and explain its cost. · Why should a trace id never be a Loki label? · How do you get from a log line to the trace of that request?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.