Why this matters
The machine rebooted at 04:12 and nobody knows why. The answer is in the log of the boot before this one - if the log survived. This lesson is how to find things in the journal fast, and how to make sure it remembers across reboots.
What you need to know already: 2.5 (journalctl -u -f), 2.9 (reading status output).
The journal is structured, not a text file
The journal is kept by the systemd-journald daemon. Every entry is a record with named fields: the message, a priority (how serious), a timestamp, the unit, the PID, the UID (user ID number of the process's user), the boot ID (a random ID for each boot), the program. journalctl is a search tool over those fields, not just a way to print a text file.
journalctl -u nginx one unit
journalctl -u nginx -f follow it
journalctl -u nginx -n 100 last 100 lines
journalctl -u nginx --since "15 min ago" --until "5 min ago"
journalctl -p err -b errors and worse, this boot
journalctl -k kernel messages only (same as dmesg)
journalctl -g 'timed out' search message text for 'timed out'
journalctl _PID=1234 any process, by field
journalctl _SYSTEMD_UNIT=ssh.service _UID=0
journalctl -o json-pretty every field, for scripting
journalctl -o cat just the messages, no metadata
Flags in that list: -n number of lines, --since/--until a time window (plain English like "15 min ago" works), -p priority, -b this boot, -k kernel (the kernel's own messages, which the dmesg command also shows), -g search ("grep") the message text, FIELD=value match a field exactly, -o output format (json-pretty shows every field of each record; cat only the message).
A normal output line reads: Sep 22 17:43:01 oncall-lab sshd[700]: Server listening on 0.0.0.0 port 22. - timestamp, hostname, program name with its PID in brackets, message.
The eight priorities
Each record has a priority (also called log level), from most to least serious:
0 emerg 1 alert 2 crit 3 err
4 warning 5 notice 6 info 7 debug
-p err means 0 through 3, not "only err" - priority filters are "this level and worse". -p warning on a noisy box is usually the right first look.
-x adds catalog text: for well-known messages, a paragraph explaining what the message means and what to do. -e jumps to the end. Try journalctl -xe after a failed unit.
Boots
journalctl --list-boots
journalctl -b this boot
journalctl -b -1 the previous boot
--list-boots prints one row per boot the journal remembers: an index (0 = this boot, -1 = the one before, ...), the boot ID, and the first and last timestamp. -b -1 is how you find out why a machine rebooted: whatever happened is in the previous boot's log; the current boot's log starts after the fact.
...unless the journal is volatile
$ ls /var/log/journal
ls: cannot access '/var/log/journal': No such file or directory
On your VM: Ubuntu Server ships
/var/log/journal, so your real journal is persistent from day one. The lab box starts volatile on purpose (simulator), so you see the failure mode and fix it once - on other distros and minimal images it is real.
If that directory does not exist, journald keeps everything in /run/log/journal - which lives on tmpfs, a filesystem held only in RAM. Every reboot wipes it. This is called volatile storage (lost at power off), as opposed to persistent storage (kept on disk). journalctl -b -1 finds nothing, and the reason your machine rebooted is gone forever.
Storage= in /etc/systemd/journald.conf (journald's own settings file):
auto (default) persistent IF /var/log/journal exists, else volatile
persistent create the directory and always persist
volatile memory only
none discard everything
Ubuntu ships auto, and a minimal image may not have the directory. Making it persistent is one command:
sudo mkdir -p /var/log/journal
sudo systemctl restart systemd-journald # or: sudo journalctl --flush
Keeping it from eating the disk
journalctl --disk-usage
sudo journalctl --vacuum-time=7d
sudo journalctl --vacuum-size=200M
--vacuum-time=7d deletes archived journal files older than 7 days; --vacuum-size=200M deletes the oldest until the total is under 200M. Or set SystemMaxUse= in journald.conf. By default the journal takes up to 10% of the filesystem, capped at 4G - which on a small disk is exactly the kind of thing you discover during an incident.
What you can now do
- filter the journal by unit, time, priority, boot and field
- read the previous boot's log to explain a reboot
- make the journal persistent and keep its size in check