OnCallReady

Lesson 5.8 · Memory & OOM · 17 min read

Reading an OOM report line by line

In plain words

Think of an accident report after a traffic jam. It has a fixed layout: who was trying to get through when it jammed, how bad the jam was (were the side roads full too?), a list of every car on the road with its size, why one car was towed, and a note confirming the road was cleared.

The kernel's OOM report is the same, in the same order every time. The first line says who invoked the OOM killer. The swap and RAM lines say how full everything was. The Tasks state table lists every candidate with rss, swapents and oom_score_adj, in pages. The oom-kill: line says whether it was the whole machine (global_oom) or one fenced area (CONSTRAINT_MEMCG) and who was chosen. The "Killed process" and oom_reaper lines confirm the tow.

The whole report

After an OOM kill most people read one line - "Killed process" - and blame whatever died. The full report the kernel writes says who asked, how full the box was, who else was there and why that one was chosen. This lesson reads it end to end.

What you need to know already: 5.7 (the OOM killer, oom_score_adj, host vs cgroup OOM), 5.1 (pages, RSS, anonymous vs file-backed).

When the OOM killer runs, the kernel writes one block to its ring buffer. Here is a host OOM on a 6 GiB box, as sudo dmesg -T shows it (trimmed where marked):

[Wed Sep 23 02:41:07 2026] python3 invoked oom-killer: gfp_mask=0x140cca(GFP_HIGHUSER_MOVABLE|__GFP_COMP), order=0, oom_score_adj=0
[Wed Sep 23 02:41:07 2026] CPU: 1 UID: 1001 PID: 5820 Comm: python3 Not tainted 7.0.0-31-generic #31-Ubuntu
[Wed Sep 23 02:41:07 2026] Hardware name: QEMU QEMU Virtual Machine, BIOS edk2-stable202408-prebuilt.qemu.org 08/13/2024
[Wed Sep 23 02:41:07 2026] Call trace:
[Wed Sep 23 02:41:07 2026]  dump_header+0x4c/0x230
[Wed Sep 23 02:41:07 2026]  oom_kill_process+0x12c/0x218
[Wed Sep 23 02:41:07 2026]  out_of_memory+0xf4/0x378
                            ... (allocator frames, then Mem-Info) ...
[Wed Sep 23 02:41:07 2026] Free swap  = 0kB
[Wed Sep 23 02:41:07 2026] Total swap = 4194300kB
[Wed Sep 23 02:41:07 2026] 1516818 pages RAM
[Wed Sep 23 02:41:07 2026] Tasks state (memory values in pages):
[Wed Sep 23 02:41:07 2026] [  pid  ]   uid  tgid total_vm      rss rss_anon rss_file rss_shmem pgtables_bytes swapents oom_score_adj name
[Wed Sep 23 02:41:07 2026] [    700]     0   700     3940     2210      540     1670         0    61440        0         -1000 sshd
[Wed Sep 23 02:41:07 2026] [   1210]  1001  1210  1050000   153000   107100    42840      3060  1224704    18420             0 java
[Wed Sep 23 02:41:07 2026] [   5820]  1001  5820   981000   812400   811100     1300         0  6500352   402100             0 python3
[Wed Sep 23 02:41:07 2026] oom-kill:constraint=CONSTRAINT_NONE,nodemask=(null),cpuset=/,mems_allowed=0,global_oom,task_memcg=/system.slice/cron.service,task=python3,pid=5820,uid=1001
[Wed Sep 23 02:41:07 2026] Out of memory: Killed process 5820 (python3) total-vm:3924000kB, anon-rss:3244400kB, file-rss:5200kB, shmem-rss:0kB, UID:1001 pgtables:6348kB oom_score_adj:0
[Wed Sep 23 02:41:07 2026] oom_reaper: reaped process 5820 (python3), now anon-rss:0kB, file-rss:0kB, shmem-rss:0kB

It looks like noise. It is five answers in a fixed order. (The CPU:, Hardware name: and Call trace: lines say where it happened inside the kernel - useful to kernel developers; skip them.)

1. Who asked - the first line

python3 invoked oom-killer names the process whose allocation failed and triggered the killer - the invoker. It is not necessarily the one that dies, and it is not necessarily the one that is leaking - it is simply the one that happened to ask for a page at the moment there were none.

order=0 means it wanted 2^0 = one single 4 KiB page: the box was completely out, not just short of big contiguous blocks.

gfp_mask ("get free pages" flags) says what kind of allocation it was. GFP_HIGHUSER_MOVABLE is ordinary process memory; GFP_KERNEL would be the kernel allocating for itself.

2. How bad - the swap and RAM lines

Free swap = 0kB with Total swap = 4194300kB: swap was full too. Every byte of RAM and swap was in use. (On a box with no swap you see Total swap = 0kB, and the OOM arrives much sooner.) 1516818 pages RAM x 4 KiB = 5.8 GiB.

3. Who was there - the Tasks state table

One row per process that was a candidate, in pages (4 KiB each on this box - multiply by 4 for KiB, divide by 256 for MiB):

[  pid  ]   uid  tgid total_vm      rss rss_anon rss_file rss_shmem pgtables_bytes swapents oom_score_adj name
[   1210]  1001  1210  1050000   153000   107100    42840      3060  1224704    18420             0 java

For a cgroup OOM this table lists only the tasks inside that cgroup - the rest of the machine was never considered.

4. Why that one - the oom-kill line

oom-kill:constraint=CONSTRAINT_NONE,...,global_oom,task_memcg=/system.slice/cron.service,task=python3,pid=5820,uid=1001

The victim is the eligible task with the highest badness - the kernel's name for the score:

badness = rss + swapents + pgtables_bytes/4096      (all in pages)
        + oom_score_adj x totalpages / 1000

where totalpages is RAM + swap for a host OOM, or the cgroup's limit for a cgroup OOM. So oom_score_adj=500 adds half of all memory to a process's score; -1000 makes it ineligible. /proc/<pid>/oom_score shows the same thing normalised to 0-1000 (plus adj), live, for every process right now.

python3: 812400 + 402100 + 1587 = 1216087 pages. java: 153000 + 18420 + 299 =

  1. python3 wins by a mile - here the victim and the culprit are the same.

They often are not.

5. What it cost - the Killed line and the reaper

Out of memory: Killed process 5820 (python3) total-vm:3924000kB, anon-rss:3244400kB, file-rss:5200kB, shmem-rss:0kB, UID:1001 pgtables:6348kB oom_score_adj:0

Same numbers, now in kB: anon-rss 3244400 kB = 3.1 GiB of heap the process held. For a cgroup OOM the line starts Memory cgroup out of memory: instead of Out of memory: - the fastest way to tell the two apart with grep.

oom_reaper: reaped process 5820 ... now anon-rss:0kB confirms the memory was actually returned. The oom_reaper is a kernel helper that frees a victim's memory right away: a victim stuck in D state (3.3) cannot exit, and the reaper frees its memory anyway so the box recovers.

Where else it shows up

journalctl -g PATTERN keeps only matching messages; -b -1 means the previous boot (2.30):

journalctl -k -g 'Killed process'           kernel log, this boot
journalctl -k -b -1 -g 'invoked oom-killer' the previous boot (persistent journal)
journalctl -u payments | grep -i oom        systemd's side of it

systemd writes its own lines for a unit whose process was killed:

payments.service: A process of this unit has been killed by the OOM killer.
payments.service: Main process exited, code=killed, status=9/KILL
payments.service: Failed with result 'oom-kill'.

status=9/KILL is signal 9 - the 137 of lesson 5.11 before the +128.

Later (Ch 16): on a Kubernetes machine, task_memcg is a long kubepods.slice/... path whose IDs tell you which container died, and Kubernetes records Reason: OOMKilled, Exit Code: 137 on it - but only for a cgroup OOM. A host OOM looks different there, which is why you still read the kernel log to be sure.

What you can now do

Why it helps

An OOM report is the most precise evidence you get in a memory incident, and most engineers scroll past it. Reading it tells you in one minute whether the machine or only one service's cgroup ran out, which cgroup the victim lived in (a cron job, a service, a systemd-run scope), whether swap was exhausted, and whether the victim was actually the big consumer. That determines the fix: raise a limit, find a leak elsewhere, or add capacity.

The task_memcg path maps a kill to a specific service, which matters because the victim's own log usually says nothing. Post-mortems and on-call handovers improve a lot when someone pastes the right three lines and explains them.

Commands in this lesson

dmesg

FAQ

Why are the numbers in the task table so small?

They are in pages, not kilobytes. On x86-64 and most arm64 systems a page is 4 KiB, so multiply by 4 for KiB or divide by 256 for MiB: 153000 pages is about 598 MiB. The exception is pgtables_bytes, which is in bytes. The final "Killed process" line uses kB, which is why the same process looks different in the two places.

How can I tell a cgroup OOM from a host OOM quickly?

Look at the oom-kill: line: constraint=CONSTRAINT_MEMCG with oom_memcg= means a cgroup limit, CONSTRAINT_NONE with global_oom means the whole machine. With grep, the summary line starts with "Memory cgroup out of memory:" for a cgroup OOM and "Out of memory:" for a host OOM. For a cgroup OOM the task table only lists that cgroup's tasks.

What does order=0 in the first line mean?

The allocation that failed asked for 2^0 = 1 contiguous page. An order-0 failure means the system genuinely had no reclaimable page at all. Higher orders (order=3 or more) mean a request for contiguous blocks, where the failure can come from fragmentation even with free memory available; those point to kernel or driver allocations rather than plain process memory.

Which process should I blame?

Not automatically the invoker or the victim. Compare the table: whose rss and swapents are large and unexpected, and whose memory grew over time. The victim is simply the highest badness. If the victim is a large, stable service and a smaller process was growing, the smaller one is the likely culprit. task_memcg tells you which service each belongs to.

What does the oom_reaper line mean?

After the kill, a kernel helper called the oom_reaper frees the victim's anonymous memory straight away, without waiting for the process to finish exiting. "reaped process ... now anon-rss:0kB" confirms the memory really came back. It matters when the victim is stuck, for example in D state waiting on storage, and could otherwise hold its memory long after being killed.

In an interview Junior

How does the OOM killer choose which process to kill?

It kills the eligible process with the highest badness, the kernel's score:

badness = rss + swapents + pgtables_bytes/4096   (pages)
        + oom_score_adj x totalpages / 1000

So it is mostly size: resident pages plus swapped pages. totalpages is RAM + swap for a host OOM, or the cgroup's limit for a cgroup OOM - where only the processes inside that cgroup are candidates. oom_score_adj is the thumb on the scale: +500 adds half of all memory to the score, -1000 makes a process ineligible (sshd in the report). /proc/PID/oom_score shows the live, normalised version for every process.

The consequence: the biggest process dies, not the guilty one. In the report, read the first line (the invoker - who asked for a page), the Tasks table (who was there, in pages), the oom-kill: line (constraint and task_memcg, which maps the victim to its service) and "Killed process".

Also asked: What does "python3 invoked oom-killer" tell you, and what does it not? · How do you map an OOM kill in the kernel log back to a systemd service? · How do you tell a host OOM from a cgroup OOM in the report?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.