OnCallReady

Lesson 5.1 · Memory & OOM · 19 min read

RSS, VSZ and shared

In plain words

Imagine booking a restaurant. You can reserve the whole back room (VSZ), but only the chairs people actually sit on are in use (RSS). And if two tables share the same bread basket, each table might claim the basket as theirs when asked, so adding up every table's claims gives more baskets than exist (shared pages counted twice).

That is why a Java process with a 512 MB heap shows 4 GB of VSZ and about 600 MB of RSS: it reserved a big room, but only 600 MB of chairs are filled. Sort processes by RSS with ps -eo pid,rss,vsz,comm --sort=-rss, but do not add RSS up. PSS in /proc/PID/smaps_rollup splits the shared basket fairly, so it adds up correctly.

Three numbers, only one of which you should act on

Someone looks at ps, sees a process "using 4 GB" on a 6 GB box and wants a bigger machine. Usually the process uses a fraction of that. This lesson is about reading a process's memory numbers correctly before anyone acts on them.

What you need to know already: 3.1 (ps, and the RSS and VSZ columns), 3.14 (/proc: every process is a directory).

Pages, first

The kernel does not hand out memory byte by byte. It hands it out in pages: fixed-size chunks of 4 KiB on this box. Every number in this chapter is really a count of pages, converted to KiB or MiB for you. (KiB = 1024 bytes, MiB = 1024 KiB. Linux tools write "kB" but mean KiB.)

A process can have a page in one of two ways:

The three numbers

VSZ   virtual size: everything the process has MAPPED. Includes memory it has
      never touched, every shared library counted in full, and reserved address
      space. A Java program with a 512M heap routinely shows 4GB of VSZ. It is
      not using 4GB. VSZ is nearly useless for capacity.

RSS   resident set size: physical pages actually in RAM right now. This is the
      number to sort by - but it over-counts, see below.

SHR   the part of RSS that is shared: libc, the binary itself, shared memory.
      Twenty processes linking libc each count its pages in their own RSS, so
      adding up everyone's RSS can give you more than the machine has.

A shared library is code many programs use (like libc, the C standard library almost every program links). It is loaded into RAM once and mapped into every process that uses it - so it shows up in each of their RSS.

So: sort by RSS, but do not add RSS up.

ps

ps -eo pid,user,rss,vsz,pmem,comm --sort=-rss | head -6: list processes; -e every process; -o choose the columns (pid, owner, RSS, VSZ, %MEM, command name); --sort=-rss sort by RSS, the minus meaning biggest first; head -6 keeps the header and the top five.

$ ps -eo pid,user,rss,vsz,pmem,comm --sort=-rss | head -6
    PID USER        RSS     VSZ %MEM COMMAND
   1210 appuser  612000 4200000 10.1 java
    415 root      27100   81300  0.4 multipathd
    640 root      23800   71400  0.4 unattended-upgr
    352 root      18400   55200  0.3 systemd-journal
      1 root      12800   22400  0.2 systemd

The top line is the orders service. It is a Java program: Java programs do not run directly on the CPU, they run inside a runtime called the JVM (Java Virtual Machine), a program that loads the Java code and runs it. So the process you see is java - the JVM - with the orders code inside it.

The java line is the pattern to recognise: VSZ seven times RSS. At start-up the JVM reserves address space for everything it might ever need - the largest heap it is allowed (the area where the program keeps its data), plus a stack per thread (each thread's scratch space for function calls). And glibc, the C library, reserves up to 64 MiB of malloc arena per thread (malloc = the standard "give me memory" call). Reserved is not used.

top says the same thing with different names

top -b -n1: -b batch mode (print instead of taking over the screen), -n1 one screen and exit.

$ top -b -n1 | head -8
top - 20:00:03 up 2:17,  1 user,  load average: 0.09, 0.09, 0.05
Tasks:  45 total,   1 running,  44 sleeping,   0 stopped,   0 zombie
%Cpu(s):  0.7 us,  0.3 sy,  0.0 ni, 99.0 id,  0.0 wa,  0.0 hi,  0.0 si,  0.0 st
MiB Mem :   5925.1 total,   2093.1 free,    966.8 used,   2865.2 buff/cache
MiB Swap:   4096.0 total,   4083.7 free,     12.3 used.   4557.2 avail Mem

    PID USER      PR  NI    VIRT    RES    SHR S  %CPU  %MEM     TIME+ COMMAND
   1210 appuser   20   0 4200000 612000 171360 S   1.3  10.1   1:46.87 java

VIRT = VSZ, RES = RSS, SHR = the shared part of RES. In interactive top, M sorts by memory, e cycles the units. The MiB Mem and MiB Swap lines are the whole machine - the next lesson (5.3) takes them apart.

The same numbers from /proc

grep -E 'A|B|C' file prints the lines of a file that match any of the words separated by | (-E allows the |). pgrep -f orders.jar prints the PID of the process whose command line contains orders.jar, and $( ) pastes that PID into the path.

$ grep -E 'VmPeak|VmSize|VmHWM|VmRSS|RssAnon|RssFile|RssShmem|VmSwap' /proc/$(pgrep -f orders.jar)/status
VmPeak:	 4284000 kB      <- the largest VSZ it ever had
VmSize:	 4200000 kB      <- VSZ
VmHWM:	  642600 kB      <- "high water mark": the largest RSS it ever had
VmRSS:	  612000 kB      <- RSS
RssAnon:	  428400 kB      <- anonymous: heap and stacks. THIS is what the OOM
                            killer really cares about.
RssFile:	  171360 kB      <- file-backed: the binary, libraries, mapped files.
                            Reclaimable - the kernel can drop it and re-read.
RssShmem:	   12240 kB      <- shared memory (tmpfs, /dev/shm, SysV shm)
VmSwap:	       0 kB      <- how much of it is in swap right now

Two new words:

RssAnon is the interesting one. Anonymous memory can only be reclaimed by moving it to swap (disk space used as overflow for RAM - lesson 5.3) or by killing the process. When you ask "how much memory does this process really cost me", RssAnon is the closest single answer. (The OOM killer is lesson 5.7: the kernel's last resort when memory runs out.)

VmHWM is worth a look after an incident: it tells you how high the process went, even if it has come back down since.

PSS: the number that adds up

smaps_rollup is a /proc file with the process's memory totals, including one number status does not have:

$ grep -E '^(Rss|Pss|Pss_Anon|Pss_File|Shared_Clean|Private_Dirty):' /proc/$(pgrep -f orders.jar)/smaps_rollup
Rss:              612000 kB
Pss:              555451 kB
Pss_Anon:         428400 kB
Pss_File:         114811 kB
Shared_Clean:      94248 kB
Private_Dirty:    440640 kB

PSS (proportional set size) divides every shared page by the number of processes sharing it. A 2 MiB library page shared by four processes counts 512 KiB in each. PSS is the only per-process figure you can sum across processes and get a number that means something.

The full smaps (one block per mapping) can be megabytes for a Java process; smaps_rollup is the cheap summary. smem (a package) prints PSS per process in a table.

Why summing RSS is wrong in both directions

ps -eo rss= prints only the RSS column with no header (the =). awk '{s+=$1} END {print s/1024 " MiB"}' adds up the first column of every line and prints the total at the end (awk gets its own lesson in 7.8).

$ ps -eo rss= | awk '{s+=$1} END {print s/1024 " MiB"}'
811.508 MiB
$ free -m | head -2
               total        used        free      shared  buff/cache   available
Mem:            5925         966        2093           4        2865        4557

Here the sum is lower than used: used also includes the kernel's own memory (the slab caches, page tables - the kernel's map from each process's addresses to real RAM - and kernel stacks), which belongs to no process. On a box running forty copies of the same web worker the sum is higher than RAM, because every copy counts the shared code and libraries in full. Either way, RSS sums are not capacity numbers:

Later (Ch 15): Kubernetes judges a container by its working set: memory.current of the container's cgroup minus inactive file pages. It includes page cache the container caused, which is why a container that reads big files can show a working set well above the RSS of its only process.

What you can now do

Why it helps

Capacity discussions and incident reviews go wrong when people read the wrong column: "the Java service is using 4 GB" from VSZ, or "our workers use 12 GB" from summed RSS on a box with 8 GB. Knowing RSS, VSZ, shared, RssAnon and PSS lets you say precisely what a process costs and what would come back if it died.

It also feeds every later memory decision. RssAnon is what the OOM killer really weighs, and what cannot be dropped like cache. And VmHWM after an incident shows how high memory went even after it dropped, which is exactly what you need for choosing a MemoryMax that will not be hit on a normal busy day.

Commands in this lesson

ps top grep free

FAQ

Why is VSZ so much bigger than RSS for Java?

VSZ counts all virtual address space mapped: the full reserved maximum heap, reserved metaspace and code cache, one stack reservation per thread, glibc's per-thread malloc arenas, and every shared library in full. Most of it is reserved and never touched, so it uses no RAM. RSS only counts pages actually resident. A 4 GB VSZ with 600 MB RSS is normal for a JVM and says nothing about memory pressure.

What is the difference between RssAnon and RssFile?

RssAnon is anonymous memory: heap, stacks and other allocations with no backing file. It can only be reclaimed by swapping or killing the process. RssFile is file-backed: the executable, libraries and memory-mapped files. The kernel can drop those pages and re-read them from disk. RssAnon is the closest single number to what a process really costs.

What is PSS and when do I need it?

Proportional set size: each shared page is divided by the number of processes sharing it. If four processes share a 2 MiB library page, each gets 512 KiB. Summing PSS across processes gives a meaningful total, which summing RSS does not. Use it when many workers share code, like a web server that starts many copies of itself. /proc/PID/smaps_rollup and the smem tool show it.

Why does the sum of RSS not match used in free?

Two effects pull in opposite directions. Shared pages are counted in every process's RSS, so the sum can exceed real usage. Meanwhile kernel memory (slab caches, page tables, kernel stacks) belongs to no process and is missing from the sum. free's used includes the kernel. So neither direction is reliable; use available for capacity and PSS for totals.

What is VmHWM and when is it useful?

The "high water mark" line in /proc/PID/status: the largest RSS the process has had since it started. After a memory spike has passed, VmRSS looks normal again, but VmHWM still shows how high it went. That is the number to compare with a memory limit you are about to set. VmPeak is the same idea for VSZ, and is much less useful.

In an interview Junior

What is the difference between RSS and VSZ?

RSS over-counts shared pages: libc is in RAM once but in every process's RSS, so do not add RSS up - forty copies of a worker can "use" more than the machine has. For a group of processes, sum PSS (from /proc/PID/smaps_rollup), which divides shared pages fairly. For one process, /proc/PID/status splits RSS into RssAnon (heap and stacks, the costly part) and RssFile (droppable), and VmHWM shows its peak.

Also asked: Forty worker processes each show 300 MB RSS on an 8 GB box. How is that possible? · What is the difference between anonymous and file-backed memory? · How do you see how high a process's memory went, after the spike has passed?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.