OnCallReady

Lesson 2.32 · systemd · 19 min read

journalctl as a query language

In plain words

Imagine a huge library where every book has labels on its spine: author, date, topic, how urgent. You could walk every shelf reading each page, or you could ask the librarian "books by this author, from last Tuesday, about fires". The second way takes seconds.

The journal is that library, and each entry has labels called fields: _SYSTEMD_UNIT, _PID, PRIORITY, _BOOT_ID. journalctl is the librarian. -u orders --since "15 min ago" is "this author, this week". -g 'timed out' searches inside the pages. -o cat gives you just the text, so you can count with sort | uniq -c, and -o json hands over every label. The skill is turning your question into labels.

The problem

A service has been flapping for an hour and its journal has 40,000 lines. Scrolling through them is hopeless. You need to ask the journal a precise question - "errors from this service in the last 15 minutes" - and get back ten lines. This lesson turns questions into journalctl filters.

What you need to know already: the journal, priorities, -u, -b, -k and fields from 2.30; reading systemctl status (2.9); the pipe | from Driving the shell (Chapter 1).

A few words first

Stop scrolling, start asking

"what did orders log in the last 15 minutes?"
journalctl -u orders --since "15 min ago"

"any errors on this box since boot?"
journalctl -b -p err

"when did payments first crash today?"
journalctl -u payments --since today -g 'exited' -n 1 -r   # newest first... or:
journalctl -u payments --since today -g 'exited' | head -1  # oldest first

(orders and payments stand for any service name.) Different filters combine with AND: every condition must hold. Two -u flags combine with OR: journalctl -u nginx -u orders shows both services mixed together in time order - exactly what you want when one talks to the other.

Time windows

--since "2026-09-22 14:00"      an exact moment (local time; this box runs in UTC)
--since "14:00" --until "14:30" today, between two times
--since today / yesterday
--since "90 min ago"            relative to now
--since -2h                     also relative
-b / -b -1                      this boot / the previous boot

Always limit the window during an incident. journalctl -u orders on a busy service can be millions of lines; with --since "10 min ago" it is instant.

Priorities are ranges

Every entry has a priority (2.30): 0 emerg, 1 alert, 2 crit, 3 err, 4 warning, 5 notice, 6 info, 7 debug. Lower number = more serious. -p means "this level or worse":

-p err        0..3: emerg, alert, crit, err
-p warning    0..4
-p 4          the same, by number
-p notice..warning   exactly 4 and 5

The gotcha: whatever a service prints to its normal output (stdout) is logged at info (6), no matter what the text says. A line reading ERROR something broke is priority 6, unless the program talks to the journal directly or starts its line with a code like <3>. So -p err finds systemd's own failure lines (Failed to start ...) but often not your application's errors. Search the text for those with -g:

# ucrash is an example unit that exits 3 on every start (not on this box)
journalctl -u ucrash -p err --no-pager
Sep 22 20:00:08 oncall-lab systemd[1]: Failed to start ucrash.service - crasher.
journalctl -u ucrash -g 'not found' --no-pager | tail -2
Sep 22 20:00:07 oncall-lab crash.sh[17398]: config /etc/app/app.yaml not found
Sep 22 20:00:07 oncall-lab crash.sh[17398]: config /etc/app/app.yaml not found

Read one line left to right: date and time, host name, the identifier of who wrote it with its PID in brackets (crash.sh[17398]), then the message. -g takes a regex and searches the message only. It ignores upper/lower case unless your pattern contains an uppercase letter.

Fields

Every entry is a set of fields (NAME=value labels, 2.30). Any FIELD=value argument is a filter:

journalctl _PID=17398                 one process (even after the unit is gone)
journalctl _SYSTEMD_UNIT=orders.service   lines the service itself wrote
journalctl UNIT=orders.service            lines systemd wrote ABOUT the service
journalctl SYSLOG_IDENTIFIER=sshd     = journalctl -t sshd (by identifier)
journalctl _UID=1001                  everything user 1001 (appuser) logged
journalctl _TRANSPORT=kernel          = -k

A UID is a user's ID number; id appuser shows it. -u orders is the first two unit filters at once, which is usually what you want. To see which fields an entry has, print one in JSON (a text format of "name": value pairs, the one web APIs use):

journalctl -u ucrash -o json-pretty -n 1 --no-pager
{
	"__REALTIME_TIMESTAMP": "1790107208900000",
	"_BOOT_ID": "32450080e3ccb68f362008f5ed686337",
	"PRIORITY": "3",
	"SYSLOG_IDENTIFIER": "systemd",
	"MESSAGE": "Failed to start ucrash.service - crasher.",
	"_PID": "1",
	"UNIT": "ucrash.service",
	...
}

__REALTIME_TIMESTAMP is the time in microseconds since 1970; _BOOT_ID says which boot it belongs to.

Output formats

-o changes how each entry is printed:

-o short          default: "Sep 22 20:00:07 host ident[pid]: message"
-o short-iso      2026-09-22T20:00:07+00:00 - sortable and unambiguous; use it in tickets
-o short-precise  with microseconds, to order events that happen very fast
-o cat            the message only - for feeding into other tools
-o json           one JSON object per line - for programs
-o verbose        every field of every entry

Counting and ranking

journalctl -u ucrash --no-pager | grep -c 'Scheduled restart job'
5
journalctl -u ucrash -o cat --no-pager | sort | uniq -c | sort -rn | head -3
      5 config /etc/app/app.yaml not found
      5 ucrash.service: Main process exited, code=exited, status=3/NOTIMPLEMENTED
      5 ucrash.service: Failed with result 'exit-code'.

The first counts restarts: systemd writes one "Scheduled restart job" line per automatic restart. The second is a ranking: -o cat drops time and PID so identical messages become identical lines, sort puts them next to each other, uniq -c counts each group, sort -rn puts the biggest count first, head -3 keeps the top three. It is the fastest way to see what a service says most.

Direction and size

-n 50     only the last 50 entries (still printed oldest first)
-r        newest first
-e        jump to the end when the pager opens
-f        follow: keep printing new lines as they arrive (Ctrl+C to stop)
--no-pager   print straight to the terminal (scripts, and anything you pipe)

The incident sequence

systemctl status X                              state, last lines
journalctl -u X --since "30 min ago" --no-pager   the story, limited in time
journalctl -u X -g 'error|exception|fatal'        the application's own complaint
journalctl -b -p warning --since "30 min ago"     what ELSE was going wrong
journalctl -k --since "30 min ago"                the kernel: memory kills, disk errors

The fourth line is the one people skip, and it is often the answer: the service failed because something else did.

What you can now do

Why it helps

During an incident you do not have time to scroll through millions of lines. Being able to say "show me everything payments and orders logged between 14:00 and 14:10, mixed in time order" or "count which error this service logs most" in one command is what makes you useful on an incident call. Limiting the time window is also what keeps journalctl fast on a busy server.

The priority gotcha matters in practice: an application's errors printed as text are priority info, so -p err shows systemd's "Failed to start" lines but not the application's own complaint. Knowing to search the text with -g instead avoids the false conclusion "there were no errors". The same habit - precise filters, a limited time window - works in every log tool you will ever use.

Commands in this lesson

journalctl

FAQ

What is the difference between _SYSTEMD_UNIT= and UNIT=?

_SYSTEMD_UNIT=orders.service matches entries written by processes inside that unit: the service's own output. UNIT=orders.service matches entries systemd itself wrote about the unit, like "Started" and "Failed with result". journalctl -u orders shows both, which is usually what you want. Fields starting with an underscore are filled in by journald, so a program cannot fake them.

How do multiple filters combine?

Different fields are combined with AND: _SYSTEMD_UNIT=x.service _PID=12 needs both to match. The same field given twice is OR: -u nginx -u orders shows both units mixed in time order. --since, --until, -p and -b narrow everything further. That is why a couple of flags can cut millions of lines down to one screen.

Is -g case-sensitive?

It is "smart": if your pattern is all lowercase, upper and lower case both match; if the pattern contains an uppercase letter, the match becomes exact. --case-sensitive=true or false forces it. The pattern is a regex matched against the message text only, so it will not find words that appear in other fields. Combine it with -u and --since to keep it fast.

Why does -n 50 still print oldest first?

-n 50 picks the last 50 matching entries but keeps them in time order, so the newest is at the bottom, next to your prompt. -r reverses the output to newest first, which is handy with -n 1 to get just the single most recent entry. Combining -r and -n 1 with -g answers "when did this last happen?".

When would I use -o json instead of -o cat?

-o cat prints only the message text - perfect for searching, sorting and counting with grep, sort and uniq -c. -o json prints each entry as one JSON object with every field, for when you need the fields themselves or want a program to read the output. -o short-iso gives unambiguous timestamps for tickets, and -o verbose shows all fields in a readable layout.

In an interview Junior

A service failed at around 14:05. How do you find the cause in the logs?

I ask the journal precise questions instead of scrolling, always with a time window:

  1. systemctl status X - state, exit code, last lines.
  2. journalctl -u X --since "13:55" --until "14:10" --no-pager - the story around the failure.
  3. journalctl -u X --since "13:55" -g 'error|exception|fatal' - the program's own complaint. -p err would miss it: whatever a service prints is logged at info.
  4. journalctl -b -p warning --since "13:55" - what else on the box was going wrong; often the service failed because something else did.
  5. journalctl -k --since "13:55" - the kernel: memory kills, disk errors.

To see what a noisy service says most: journalctl -u X -o cat | sort | uniq -c | sort -rn | head. For a ticket, -o short-iso gives sortable, unambiguous timestamps.

Also asked: What does "structured logging" mean in the journal, and why does it matter? · What is the difference between _SYSTEMD_UNIT= and UNIT= matches? · How do you show two services' logs interleaved in time order?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.