OnCallReady

Lesson 5.7 · Memory & OOM · 16 min read

The OOM killer

In plain words

Imagine a lifeboat that can hold only so much weight. When it is about to sink and nobody can bail out more water, the captain has to push someone overboard. The rule is simple: the heaviest passenger goes first, unless they have a special badge. The heaviest is not necessarily the one who overloaded the boat; it is just the biggest.

The OOM killer is that captain. When the kernel cannot find a free page even after dropping cache and using swap, it gives each process a score (/proc/PID/oom_score, based on its memory share) and kills the highest with SIGKILL. oom_score_adj is the badge: -1000 means never me, +1000 means me first. journalctl -k | grep -i oom is the ship's log that tells you who went overboard and why.

What actually happens

A service dies at 2am with no error in its own log. Often the kernel killed it because the machine ran out of memory - and picked it by size, not by guilt. This lesson is how the kernel chooses, where it writes it down, and how to protect what you need.

What you need to know already: 5.1 (RSS, anonymous memory), 5.3 (page cache, swap, overcommit), 3.6 (signals, SIGKILL), 2.30 (journald, journalctl -k).

OOM means "out of memory". When the kernel needs a page and cannot reclaim one - the page cache is gone, swap is full or absent - it does not fail the allocation. It picks a process and kills it with SIGKILL, which frees that process's memory at once. The part of the kernel that chooses is the OOM killer.

The score

Every process gets a score, roughly "proportion of RAM + swap this process is using", 0-1000, plus an adjustment:

cat /proc/<pid>/oom_score        the computed score - biggest gets killed
cat /proc/<pid>/oom_score_adj    -1000 .. +1000, your thumb on the scale

oom_score_adj = -1000 makes a process effectively immune. +1000 makes it the first to go. Anything in between shifts the score by that many thousandths of total memory: +500 counts as if the process held an extra half of the machine.

The line below is a bash for loop (loops get their own lesson in Ch 6): for each PID in the list, it prints the PID, the command name (/proc/<pid>/comm), the score and the adjustment.

$ for p in $(pgrep -f orders.jar) $(pgrep -o sshd) 1; do echo "$p $(cat /proc/$p/comm) $(cat /proc/$p/oom_score) $(cat /proc/$p/oom_score_adj)"; done
1210 java 60 0
700 sshd 1 0
1 systemd 1 0

Columns: PID, name, oom_score, oom_score_adj. java holds 612 MB out of RAM + swap's 10 GB total: 60 thousandths. Everything else is a rounding error - so on this box, as things stand, java dies first.

The consequence people trip over: the biggest process is not necessarily the guilty one. A leaking 200MB script can push the box over the edge, and the kernel kills your 4GB database because it scores highest. The kill is about size, not blame.

Finding the evidence

The OOM killer is part of the kernel, so it writes to the kernel ring buffer: a fixed-size log in the kernel's memory, where the newest messages overwrite the oldest. Two ways to read it:

journalctl -k | grep -i oom
sudo dmesg -T | grep -i -E 'oom|killed process'

(grep -i ignores upper/lower case.)

python3 invoked oom-killer: gfp_mask=0x140dca(...), order=0, oom_score_adj=0
oom-kill:constraint=CONSTRAINT_NONE, ..., global_oom, task=python3, pid=2211
Out of memory: Killed process 2211 (python3) total-vm:5410244kB,
               anon-rss:5102336kB, ...

Two fields to read carefully:

Those are two completely different incidents, and the log line is what distinguishes them.

Note also: "python3 invoked oom-killer" names the process that requested the page, not necessarily the one killed. Read down to "Killed process". The next lesson takes a full report apart line by line.

dmesg needs root on Ubuntu (kernel.dmesg_restrict = 1). journalctl -k works without it for members of adm, and with a persistent journal (2.30) it survives a reboot - journalctl -k -b -1 is how you find the OOM that preceded a crash.

Protecting what you need to get back in

echo -1000 | sudo tee /proc/$(pgrep -o sshd)/oom_score_adj

pgrep -o picks the oldest match - the listening sshd, not your session's child. Now sshd will not be chosen, and you can still log in to a box that is thrashing. It lasts until sshd restarts; do it permanently in the unit:

[Service]
OOMScoreAdjust=-1000

(sudo systemctl edit ssh, then restart it - drop-ins, 2.3.) systemd already does this for its own critical services. Worth doing for sshd on anything you cannot walk up to.

Careful with -1000 on anything else. A service that can never be chosen, and that leaks, leaves the kernel killing everything around it - or, with nothing else left, panicking (the kernel stops the whole machine). For an important service, -500 says "prefer others" without saying "never".

Later (Ch 16): Kubernetes sets oom_score_adj on every container for you, from how its memory was requested: containers that reserved exactly what they may use get -997 (killed last), containers that reserved nothing get 1000 (killed first), everything else in between.

Userspace killers

systemd-oomd is a daemon that watches memory pressure - how much of the time processes are stalled waiting for memory, reported by the kernel as PSI (pressure stall information) in /proc/pressure/memory - and kills a whole cgroup before the kernel has to. If it is running, a kill can come from it rather than the kernel - it logs to its own unit (journalctl -u systemd-oomd), not to the kernel log. Check with systemctl status systemd-oomd.

What you can now do

Why it helps

An OOM kill is often the root cause behind "the database just died" or "a random service restarted at 2am". Being able to find it in the kernel log, and read whether it was a global or a cgroup OOM, turns a mysterious crash into a precise finding. The biggest-not-guilty rule is crucial in post-mortems: the killed process may be a victim of a small leaking cron job.

Protecting sshd with OOMScoreAdjust=-1000 keeps you able to log in to a struggling box. Choosing negative adjustments for important services, and limits for risky ones, is literally choosing who dies first when the machine runs out. Userspace killers like systemd-oomd add a second place to look.

Commands in this lesson

cat journalctl

FAQ

Why did the OOM killer kill my database instead of the process that leaked?

The kernel chooses by badness, essentially memory footprint (RSS, swap and page tables) plus oom_score_adj, not by who allocated most recently or who grew fastest. A 4 GB database outweighs a 200 MB leaking script, even if the script caused the shortage. Protect critical services with a negative OOMScoreAdjust, and put limits on the others so they hit their own cgroup limit first.

What does "X invoked oom-killer" mean?

It names the process whose allocation could not be satisfied and which therefore triggered the OOM killer. It is often not the victim; the victim appears later in the "Killed process" line. The invoker is just whoever asked for a page at the moment memory ran out, so it is weak evidence of guilt. Read the whole report.

Is setting oom_score_adj to -1000 a good idea?

For sshd and a few essential agents, yes, so you can still reach and repair the box. For large application services, be careful: if an immune service leaks, the kernel kills everything else around it, and with nothing left it may panic or stall. A value like -500 expresses "prefer other victims" without immunity. Better still, give services memory limits.

Why do I need root for dmesg?

Ubuntu sets kernel.dmesg_restrict=1, so reading the kernel ring buffer requires CAP_SYSLOG. Kernel messages can reveal addresses and other sensitive details. Use sudo dmesg -T, where -T converts timestamps to wall-clock time, or journalctl -k, which reads the same messages from the journal and works across boots if the journal is persistent. Members of adm can often read the journal.

What is systemd-oomd and how is it different?

A userspace daemon that monitors memory pressure (PSI) and swap use per cgroup, and kills a whole cgroup before the kernel is forced into an OOM. It acts earlier and on policy, which can keep a desktop or server responsive. Its kills are logged in journalctl -u systemd-oomd, not in the kernel log, which surprises people. Ubuntu enables it on desktops; check systemctl status systemd-oomd on servers.

In an interview Junior

How do you find out whether a process was killed by the OOM killer?

The OOM killer is part of the kernel, so the evidence is in the kernel log:

The kernel picks by size (/proc/PID/oom_score), not by guilt.

Also asked: How does the OOM killer choose which process to kill? · How would you protect sshd from the OOM killer? · Why is oom_score_adj=-1000 on an ordinary service risky?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.