OnCallReady

Lesson 5.3 · Memory & OOM · 21 min read

free, available, and the page cache

In plain words

Imagine a fridge. Some shelves are completely empty (free). Some hold leftovers you might eat again soon (cache). If you bring home new groceries, you just throw out some leftovers and use the space. So the question "how much can I still put in?" is empty shelves plus throw-out-able leftovers. That is available.

free -m shows all of it. People panic at a small free column, but a Linux fridge full of leftovers is a well-used fridge. The only real worry is when available runs low and you start moving food to the garage and back all day (swap in and out, si/so in vmstat 1). Dropping caches just throws away leftovers you might have wanted; it does not make the fridge bigger.

The column everyone reads is the wrong one

A dashboard says "memory 90% used" and someone wants to resize the box. Most of the time the box is fine: Linux fills idle RAM with a cache on purpose. This lesson is how to tell a full cache from a real memory shortage.

What you need to know already: 5.1 (pages, RSS, anonymous memory), 1.11 (sudo, and why | sudo tee beats sudo echo >), 3.5 (vmstat).

The page cache

When a program reads a file, the kernel keeps a copy of those pages in RAM, in case anyone reads them again - disk is thousands of times slower than RAM. That copy is the page cache. It grows into whatever RAM nobody else is using, and the kernel throws pages out of it the moment a program needs the memory.

free -h prints the machine's memory in human units (-h: Gi, Mi):

$ free -h
               total        used        free      shared  buff/cache   available
Mem:           5.8Gi      967Mi      2.0Gi      4.1Mi       2.8Gi       4.5Gi
Swap:          4.0Gi       12Mi      4.0Gi

So "the server only has 2GB free!" on a box with 4.5GB available is not a problem. It is Linux doing its job. An idle Linux box with lots of free memory has simply not read anything yet.

Alert on available, never on free.

The flags you will use

$ free -m
               total        used        free      shared  buff/cache   available
Mem:            5925         966        2093           4        2865        4557
Swap:           4095          12        4083
$ free -h -w
               total        used        free      shared     buffers       cache   available
Mem:           5.8Gi      967Mi      2.0Gi      4.1Mi        86Mi       2.7Gi       4.5Gi

-m MiB (good for scripts, whole numbers), -h human units, -w wide: splits buff/cache into buffers (the disk's own bookkeeping - filesystem metadata) and cache (file contents). free -s 2 repeats every two seconds. The plain free default is KiB.

free is /proc/meminfo in columns

/proc/meminfo is the kernel's memory report, one Name: value kB per line. free just reads it:

$ grep -E 'MemTotal|MemFree|MemAvailable|Buffers|^Cached|SwapTotal|SwapFree|AnonPages|Shmem|Slab|SReclaimable' /proc/meminfo
MemTotal:        6067272 kB
MemFree:         2143309 kB
MemAvailable:    4666549 kB
Buffers:           88020 kB
Cached:          2640600 kB
SwapTotal:       4194300 kB
SwapFree:        4181756 kB
AnonPages:        791970 kB
Shmem:              4212 kB
Slab:             205380 kB
SReclaimable:     146700 kB

Later (Ch 10): Inside a container, /proc/meminfo still shows the host's memory (unless the runtime uses lxcfs). That is why free inside a container lies, and why programs have to read their cgroup's limit instead.

vmstat: is it swapping right now?

You met vmstat in 3.5 for the r and b columns. vmstat 1 5 prints a line every 1 second, 5 times. Now read its memory and swap columns:

$ vmstat 1 5
procs -----------memory---------- ---swap-- -----io---- -system-- -------cpu-------
 r  b   swpd   free   buff  cache   si   so    bi    bo   in   cs us sy id wa st gu
 1  0  12544 2143309  88020 2845980    0    0    31    12  128  201  1  1 98  0  0  0
 1  0  12544 2143309  88020 2845980    0    0     0     4  197  331  1  1 98  0  0  0
 1  0  12544 2143309  88020 2845980    0    0     0     4  204  342  1  1 98  0  0  0

Ignore the first line - it is the average since boot. Then:

Proving the cache is available

sync                                    flush dirty pages first
echo 3 | sudo tee /proc/sys/vm/drop_caches
free -h

Dirty pages are cached pages that were changed in RAM and not yet written back to disk. sync writes them out, so they become clean and droppable. Writing a number into /proc/sys/vm/drop_caches tells the kernel to throw caches away.

buff/cache collapses, free shoots up - and available barely moves, because that memory was already counted as available. You have not gained anything; you have thrown away a cache that now has to be re-read from disk.

1 drops the page cache, 2 dentries and inodes, 3 both. Note the sudo tee (1.11): sudo echo 3 > /proc/sys/vm/drop_caches fails, because the redirection is done by your unprivileged shell. Never do this in production as a "fix". It is a demonstration.

Swap

Swap is disk used as overflow for RAM: the kernel moves anonymous pages nobody has used lately out to disk, and reads them back if they are needed. It does not add memory, it adds a slow tier.

swapon --show lists the swap areas:

$ swapon --show
NAME      TYPE SIZE  USED PRIO
/swap.img file   4G 12.3M   -2

NAME the file or partition, TYPE file or partition, SIZE and USED, PRIO which swap area is used first. (/swap.img is the same file you met in the 2.36 incident.)

vm.swappiness (0-200 on current kernels, default 60) is how eagerly the kernel swaps anonymous pages out rather than dropping page cache. Lower means "keep processes in RAM, drop cache first". 10 is a common setting for a database.

sysctl reads and sets kernel settings:

$ sysctl vm.swappiness
vm.swappiness = 60
$ sudo sysctl -w vm.swappiness=10
vm.swappiness = 10

sysctl -w (write) lasts until reboot. Permanent: a file in /etc/sysctl.d/, e.g. echo 'vm.swappiness = 10' | sudo tee /etc/sysctl.d/60-swappiness.conf. sysctl is only a front end: vm.swappiness is the file /proc/sys/vm/swappiness.

A little swap in use is fine - often it is pages touched once at boot and never again. Swap churn (constant si/so in vmstat) is the pathology: the box grinds and everything crawls, which is usually worse than a fast OOM kill.

This is also why the kubelet in the 2.36 incident refused to start with swap on: it was written to assume memory limits are real RAM.

Later (Ch 15): Kubernetes required swap off on its machines for years, because its whole model of memory limits assumes RAM. Newer releases support swap, off by default and opt-in per machine.

Overcommit: why malloc succeeds and the OOM comes later

$ sysctl vm.overcommit_memory
vm.overcommit_memory = 0
$ grep -E 'CommitLimit|Committed_AS' /proc/meminfo
CommitLimit:     7227936 kB
Committed_AS:    2078922 kB

Linux hands out address space it does not have: that is overcommit. malloc(4 GB) on a 6 GB box with 2 GB free succeeds; pages are only found when the program touches them. Committed_AS is how much has been promised in total; CommitLimit is the ceiling used only in strict mode.

Mode 0 (heuristic, the default) refuses only absurd requests, 1 never refuses, 2 strict accounting against CommitLimit. With the default, "out of memory" is not an error your program gets back from malloc - it is the OOM killer arriving later, at whoever is biggest. That is the reason lesson 5.7 exists.

What you can now do

Why it helps

This is probably the single most common false alarm in operations: "memory at 95%" alerts built on used or free instead of available. Knowing the columns lets you close those tickets in a minute and fix the alert, and it is one of the Notion interview questions.

The real memory problems are distinct and recognisable: available dropping steadily (a leak), si/so sustained in vmstat (thrashing, the box crawling), shmem growing (a full tmpfs). Swap and swappiness settings come up when tuning database hosts, and software that refuses to run with swap explains the kubelet failing in the node incident you already fixed. Overcommit explains why malloc succeeds and the OOM killer arrives later.

Commands in this lesson

free grep vmstat cat swapon sysctl

FAQ

What exactly is in buff/cache?

Buffers are cached block-device metadata (filesystem structures read from the disk). Cache is the page cache: contents of files that were read or written, kept in RAM for reuse, plus tmpfs and shared memory pages, which are counted there too. Most of it is reclaimable instantly, except the shared (tmpfs, shm) part, which stays until those files are deleted. free -w shows buffers and cache separately.

Should I ever drop caches?

Not in production as a fix. echo 3 > /proc/sys/vm/drop_caches throws away clean page cache and dentry/inode caches. It makes free look larger while available barely changes, and performance drops because files must be read from disk again. It is useful for benchmarks (cold-cache measurements) and demonstrations. If a box has real memory pressure, dropping caches does not help, because the kernel already reclaims cache first.

How do I know if a box is thrashing?

vmstat 1 shows sustained non-zero si and so (pages swapped in and out per second) together, usually with high wa and slow response. Swap used (swpd) being non-zero on its own is fine; it is continuous movement that hurts. /proc/pressure/memory gives the share of time tasks stalled on memory, and some or full values of several percent mean real pressure.

What does vm.swappiness actually control?

How the kernel balances reclaiming anonymous memory (by swapping it) against reclaiming page cache (by dropping it) under pressure. Higher values make it more willing to swap anonymous pages; lower values make it prefer dropping cache. Default 60, range 0-200 on current kernels. It does not decide whether swap is used at all, and 0 does not disable swap. Databases often use 1-10.

How do I change a kernel setting like swappiness permanently?

sudo sysctl -w vm.swappiness=10 changes it right now, until the next reboot. To keep it, put the line vm.swappiness = 10 in a file under /etc/sysctl.d/, for example with echo 'vm.swappiness = 10' | sudo tee /etc/sysctl.d/60-swappiness.conf, and load it with sudo sysctl --system. Every sysctl name is also a file under /proc/sys/ with dots turned into slashes.

In an interview Junior

Why is almost no free memory on a Linux server usually not a problem?

Because Linux fills idle RAM with the page cache - copies of files it has read, kept in case they are read again - and throws those pages out the moment a program needs the memory. Unused RAM is wasted RAM.

In free -m: free is memory doing nothing (low is fine), buff/cache is mostly that cache, and available is the kernel's estimate of what a new program could get right now, cache included. That is the number to read and to alert on. 2 GB free with 4.5 GB available is a healthy box.

Signs of a real shortage: low available, and in vmstat 1 5 sustained si/so (swapping in and out - thrashing). Dropping the cache with echo 3 | sudo tee /proc/sys/vm/drop_caches proves it: free shoots up, available barely moves. Never do that as a "fix". One exception: shared (tmpfs) is counted in buff/cache but is not reclaimable.

Also asked: Explain every column of free -m. · What is swappiness, and why would you lower it? · What is memory overcommit, and why does it mean malloc rarely fails?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.