Why logs need their own database
The metrics say checkout started failing at 20:01. They cannot say why: that is in the exception the app logged. On this one box journalctl -u orders finds it. On forty servers, or with the container that logged it already gone, there is no box to ssh into. You need every log line from every machine in one place, searchable by the same labels you use in Prometheus.
What you need to know already: journald and journalctl -u (2.30, 2.32); structured JSON logs and trace ids (27.1); labels, series and cardinality (27.1, 27.2); PromQL selectors, matchers (=, !=, =~), rate(), sum by and topk (27.8); regular expressions (7.3); jq (7.11).
The words you need first
- Loki - a log database from the makers of Grafana. It stores log lines and answers queries about them. It runs on this box as
loki.service, port 3100 (27.1). - Log stream - all the lines that share exactly the same labels. A stream's labels look like a Prometheus series:
{job="systemd-journal", unit="orders.service", host="oncall-lab"}. Every line of orders.service on this host lands in that one stream. - Chunk - a compressed block of one stream's lines, the way Loki stores them on disk.
- Index - the lookup table a database keeps so it can find things without reading everything. Loki's index holds only the labels.
- LogQL - Loki's query language, built to look like PromQL.
- Log shipper - a program on each machine that reads local logs and sends them to Loki. On this box that is Grafana Alloy, configured in
/etc/alloy/config.alloy: it reads the systemd journal. (Promtail used to do this job; it was deprecated in 2025, reached end of life in early 2026, and Alloy replaced it.)
Indexed by labels, not by content
Loki indexes the labels of each stream; the text of each line is only compressed into chunks and read when you query. Consequences:
- Selecting by label is cheap: the index says which chunks to open. Searching the text reads every line of the chosen streams over the time range. So always narrow by labels first.
- Labels must have few values (low cardinality, 27.1). Never a user id, trace id, request id, or a status code with hundreds of values as a label: every distinct label set is a new stream with its own chunks, and Loki will refuse or slow to a crawl.
- It is cheap to keep, because there is no index of every word (that is what makes systems like Elasticsearch - a search engine that indexes every word of every line - fast to search but expensive to run).
LogQL, part 1: pick streams, then filter lines
Every log query starts with a stream selector in curly braces, the same syntax as a PromQL selector. Then optional line filters, which keep or drop lines by their text:
{unit="orders.service"} stream selector (required)
{unit="orders.service"} |= "ERROR" keep lines that contain ERROR
{unit="orders.service"} != "health" drop lines that contain health
{unit="orders.service"} |~ "Hikari.*timed out" keep lines matching a regex
{unit="orders.service"} !~ "DEBUG|TRACE" drop lines matching a regex
Read |= as "pipe into: contains", |~ as "contains a regex match"; the ! versions invert. Filters run left to right and are the cheapest thing you can do to content, so put the one that throws away the most lines first.
logcli is Loki's command-line client, the terminal version of Grafana's log view:
$ logcli query '{unit="orders.service"} |= "ERROR"' --limit=2
http://localhost:3100/loki/api/v1/query_range?direction=BACKWARD&end=...&limit=2&query=...
Common labels: {host="oncall-lab", job="systemd-journal", unit="orders.service"}
2026-09-22T20:02:24Z {} {"@timestamp":"2026-09-22T20:02:24.000Z","@version":"1","message":"Request processing failed: ...","level":"ERROR",...}
2026-09-22T20:02:20Z {} {"@timestamp":"2026-09-22T20:02:20.000Z",...}
logcli query '<LogQL>'- fetch matching lines. Single-quote the query so the shell leaves its double quotes and|alone.--limit=2- at most 2 lines (default 30), newest first.--since=1h- how far back to look (the default is 1 hour).--since=24hfor a day.-o raw- print only the log lines, nothing else, for piping intojq,sortor$( ).
The output: the first two lines go to stderr - the URL logcli called (Loki's HTTP API) and the common labels every result shares. Then one line per log entry: its timestamp, the labels that are not common ({} = none left over), and the line itself - here a JSON document.
LogQL, part 2: parsers turn text into labels
{unit="orders.service"} | json
A parser reads each line and turns its fields into labels - only for the rest of this one query; nothing is indexed or stored. | json parses JSON lines. Field names are cleaned up to be valid label names: @timestamp becomes _timestamp, http.uri becomes http_uri. | logfmt does the same for logfmt lines, the key=value key=value style Prometheus itself logs in.
After a parser, two more pipeline stages become useful:
- a label filter keeps lines whose parsed field has a value:
| level="ERROR". It is not a text search - a message that merely contains the word ERROR does not match. line_formatrewrites each line from a template:{{.traceId}}is the value of thetraceIdfield (the{{.name}}syntax is Go's template language).
{unit="orders.service"} | json | level="ERROR" | line_format "{{.traceId}} {{.message}}"
2026-09-22T20:02:24Z {} 4db9ca9a4eb9cc2d4bb9c7744cb9c907 Request processing failed: org.springframework.jdbc.CannotGetJdbcConnectionException: Failed to obtain JDBC Connection
Piece by piece: select the orders stream, parse every line as JSON, keep the ones whose level field is ERROR, print only the trace id and message. The output is one error per line, easy to read.
Label filters compare numbers too: | json | duration_ms > 1000, | logfmt | status >= 500. A line the parser could not read gets a special label __error__; | __error__="" keeps only the lines that parsed.
LogQL, part 3: metric queries turn logs into numbers
Wrap a log query and a time range in a range function and the result is no longer lines but numbers, exactly like a PromQL vector (27.8):
count_over_time({unit="orders.service"} |= "ERROR" [5m]) number of lines in the last 5m, per stream
rate({unit="orders.service"} |= "ERROR" [5m]) lines per second
sum by (level) (count_over_time({unit="orders.service"} | json [15m]))
topk(3, sum by (http_uri) (count_over_time({unit="orders.service"} | json | level="ERROR" [1h])))
Take the third one apart, inside out:
{unit="orders.service"} | json- every orders line, parsed; now each line has alevellabel.[15m]- look at the last 15 minutes.count_over_time(...)- count lines, one number per distinct label set.sum by (level) (...)- add them up, keeping onlylevel: one number per level.
The fourth adds topk(3, ...): keep the three biggest, i.e. the three endpoints that logged the most errors in the last hour.
logcli instant-query runs a metric query at one moment (now) and prints JSON:
$ logcli instant-query 'sum by (level) (count_over_time({unit="orders.service"} | json [15m]))'
[
{
"metric": { "level": "ERROR" },
"value": [1790107400, "38"]
},
{
"metric": { "level": "WARN" },
"value": [1790107400, "11"]
}
]
One object per result: metric holds its labels, value is [Unix time, "value as a string"] - 38 ERROR and 11 WARN lines in the last 15 minutes.
sum by (level) works on a label that came from | json because the parser created it. This is also how you alert on logs when there is no metric: Loki's ruler evaluates LogQL rules like Prometheus does and sends to the same Alertmanager (28.7). Prefer a real metric when there is one: counting log lines is more expensive and silently breaks when someone rewords a message.
From a log line to a trace
The traceId in each JSON line (27.1) is the bridge to tracing, the next lesson. In Grafana, a "derived field" on the Loki data source turns it into a link that opens the trace. Here you copy it into trace <id> (simulator). The opposite direction - from a trace to the logs of that one request - is a line filter on the id:
{unit="orders.service"} |= "4db9ca9a4eb9cc2d4bb9c7744cb9c907"
That is why the trace id belongs in the line, never in a label: a filter finds it cheaply, while a label would create a stream per request.
journalctl is still there
On one box, journalctl -u orders -o cat | jq -r 'select(.level=="ERROR") | .message' answers the same questions without Loki (-o cat prints only the message, 2.32). What Loki adds is every box and every container in one query, and retention after the container is gone: in Kubernetes the logs of a crashed pod's previous container (kubectl logs --previous) disappear the moment the pod is deleted.
What you can now do
- Write a LogQL query in the right order: stream selector, line filters, parser, label filters,
line_format. - Turn logs into numbers with
count_over_timeandsum by, and run both kinds of query withlogcli. - Explain why ids stay in the log line and never become Loki labels.