What actually happens
A service dies at 2am with no error in its own log. Often the kernel killed it because the machine ran out of memory - and picked it by size, not by guilt. This lesson is how the kernel chooses, where it writes it down, and how to protect what you need.
What you need to know already: 5.1 (RSS, anonymous memory), 5.3 (page cache, swap, overcommit), 3.6 (signals, SIGKILL), 2.30 (journald, journalctl -k).
OOM means "out of memory". When the kernel needs a page and cannot reclaim one - the page cache is gone, swap is full or absent - it does not fail the allocation. It picks a process and kills it with SIGKILL, which frees that process's memory at once. The part of the kernel that chooses is the OOM killer.
The score
Every process gets a score, roughly "proportion of RAM + swap this process is using", 0-1000, plus an adjustment:
cat /proc/<pid>/oom_score the computed score - biggest gets killed
cat /proc/<pid>/oom_score_adj -1000 .. +1000, your thumb on the scale
oom_score_adj = -1000 makes a process effectively immune. +1000 makes it the first to go. Anything in between shifts the score by that many thousandths of total memory: +500 counts as if the process held an extra half of the machine.
The line below is a bash for loop (loops get their own lesson in Ch 6): for each PID in the list, it prints the PID, the command name (/proc/<pid>/comm), the score and the adjustment.
$ for p in $(pgrep -f orders.jar) $(pgrep -o sshd) 1; do echo "$p $(cat /proc/$p/comm) $(cat /proc/$p/oom_score) $(cat /proc/$p/oom_score_adj)"; done
1210 java 60 0
700 sshd 1 0
1 systemd 1 0
Columns: PID, name, oom_score, oom_score_adj. java holds 612 MB out of RAM + swap's 10 GB total: 60 thousandths. Everything else is a rounding error - so on this box, as things stand, java dies first.
The consequence people trip over: the biggest process is not necessarily the guilty one. A leaking 200MB script can push the box over the edge, and the kernel kills your 4GB database because it scores highest. The kill is about size, not blame.
Finding the evidence
The OOM killer is part of the kernel, so it writes to the kernel ring buffer: a fixed-size log in the kernel's memory, where the newest messages overwrite the oldest. Two ways to read it:
journalctl -k- the kernel messages as journald stored them (2.30).dmesg- reads the ring buffer directly.-Tturns seconds-since-boot into wall-clock time.
journalctl -k | grep -i oom
sudo dmesg -T | grep -i -E 'oom|killed process'
(grep -i ignores upper/lower case.)
python3 invoked oom-killer: gfp_mask=0x140dca(...), order=0, oom_score_adj=0
oom-kill:constraint=CONSTRAINT_NONE, ..., global_oom, task=python3, pid=2211
Out of memory: Killed process 2211 (python3) total-vm:5410244kB,
anon-rss:5102336kB, ...
Two fields to read carefully:
constraint=CONSTRAINT_NONE+global_oom- the whole machine ran out. A host OOM.constraint=CONSTRAINT_MEMCG+oom_memcg=/system.slice/foo.service- a cgroup hit its own limit (theMemoryMax=from 2.26; "memcg" = memory cgroup). The machine was fine. A cgroup OOM - lesson 5.11.
Those are two completely different incidents, and the log line is what distinguishes them.
Note also: "python3 invoked oom-killer" names the process that requested the page, not necessarily the one killed. Read down to "Killed process". The next lesson takes a full report apart line by line.
dmesg needs root on Ubuntu (kernel.dmesg_restrict = 1). journalctl -k works without it for members of adm, and with a persistent journal (2.30) it survives a reboot - journalctl -k -b -1 is how you find the OOM that preceded a crash.
Protecting what you need to get back in
echo -1000 | sudo tee /proc/$(pgrep -o sshd)/oom_score_adj
pgrep -o picks the oldest match - the listening sshd, not your session's child. Now sshd will not be chosen, and you can still log in to a box that is thrashing. It lasts until sshd restarts; do it permanently in the unit:
[Service]
OOMScoreAdjust=-1000
(sudo systemctl edit ssh, then restart it - drop-ins, 2.3.) systemd already does this for its own critical services. Worth doing for sshd on anything you cannot walk up to.
Careful with -1000 on anything else. A service that can never be chosen, and that leaks, leaves the kernel killing everything around it - or, with nothing else left, panicking (the kernel stops the whole machine). For an important service, -500 says "prefer others" without saying "never".
Later (Ch 16): Kubernetes sets
oom_score_adjon every container for you, from how its memory was requested: containers that reserved exactly what they may use get-997(killed last), containers that reserved nothing get1000(killed first), everything else in between.
Userspace killers
systemd-oomd is a daemon that watches memory pressure - how much of the time processes are stalled waiting for memory, reported by the kernel as PSI (pressure stall information) in /proc/pressure/memory - and kills a whole cgroup before the kernel has to. If it is running, a kill can come from it rather than the kernel - it logs to its own unit (journalctl -u systemd-oomd), not to the kernel log. Check with systemctl status systemd-oomd.
What you can now do
- Find an OOM kill in the kernel log and say whether the machine or a cgroup ran out.
- Predict who the kernel would kill right now from
/proc/<pid>/oom_score. - Protect sshd with
oom_score_adj/OOMScoreAdjust=.