The column everyone reads is the wrong one
A dashboard says "memory 90% used" and someone wants to resize the box. Most of the time the box is fine: Linux fills idle RAM with a cache on purpose. This lesson is how to tell a full cache from a real memory shortage.
What you need to know already: 5.1 (pages, RSS, anonymous memory), 1.11 (sudo, and why | sudo tee beats sudo echo >), 3.5 (vmstat).
The page cache
When a program reads a file, the kernel keeps a copy of those pages in RAM, in case anyone reads them again - disk is thousands of times slower than RAM. That copy is the page cache. It grows into whatever RAM nobody else is using, and the kernel throws pages out of it the moment a program needs the memory.
free -h prints the machine's memory in human units (-h: Gi, Mi):
$ free -h
total used free shared buff/cache available
Mem: 5.8Gi 967Mi 2.0Gi 4.1Mi 2.8Gi 4.5Gi
Swap: 4.0Gi 12Mi 4.0Gi
- total - the RAM the kernel can use.
- free - memory doing absolutely nothing. Low is good: unused RAM is wasted RAM.
- buff/cache - mostly the page cache. Counted as used, but it is a cache.
- available - an estimate of how much a new process could get right now, including everything the kernel would reclaim from the cache to give it. This is the number that matters.
- used - total minus free minus buff/cache. Processes plus the kernel.
- shared - tmpfs (a filesystem that lives in RAM, like
/runand/dev/shm) and shared memory. It is part of buff/cache and it is not reclaimable - a full tmpfs is real memory use. - The Swap line - swap space on disk: total, used, free.
So "the server only has 2GB free!" on a box with 4.5GB available is not a problem. It is Linux doing its job. An idle Linux box with lots of free memory has simply not read anything yet.
Alert on available, never on free.
The flags you will use
$ free -m
total used free shared buff/cache available
Mem: 5925 966 2093 4 2865 4557
Swap: 4095 12 4083
$ free -h -w
total used free shared buffers cache available
Mem: 5.8Gi 967Mi 2.0Gi 4.1Mi 86Mi 2.7Gi 4.5Gi
-m MiB (good for scripts, whole numbers), -h human units, -w wide: splits buff/cache into buffers (the disk's own bookkeeping - filesystem metadata) and cache (file contents). free -s 2 repeats every two seconds. The plain free default is KiB.
free is /proc/meminfo in columns
/proc/meminfo is the kernel's memory report, one Name: value kB per line. free just reads it:
$ grep -E 'MemTotal|MemFree|MemAvailable|Buffers|^Cached|SwapTotal|SwapFree|AnonPages|Shmem|Slab|SReclaimable' /proc/meminfo
MemTotal: 6067272 kB
MemFree: 2143309 kB
MemAvailable: 4666549 kB
Buffers: 88020 kB
Cached: 2640600 kB
SwapTotal: 4194300 kB
SwapFree: 4181756 kB
AnonPages: 791970 kB
Shmem: 4212 kB
Slab: 205380 kB
SReclaimable: 146700 kB
MemTotal/MemFree- free's total and free.MemAvailableis free's "available" - the kernel computes it (free memory + reclaimable cache + reclaimable slab - reserves), free just prints it.Buffers+Cached- free's buff/cache.AnonPages- all anonymous memory of all processes: heaps and stacks.Shmem- free's shared.Slab/SReclaimable- the kernel's own caches of small objects, like dentries (the kernel's cache of file names it has looked up) and inodes (4.13). A box with a hugeSUnreclaimhas a kernel-side leak - rare, but it happens with some drivers and with millions of network connections.
Later (Ch 10): Inside a container,
/proc/meminfostill shows the host's memory (unless the runtime uses lxcfs). That is whyfreeinside a container lies, and why programs have to read their cgroup's limit instead.
vmstat: is it swapping right now?
You met vmstat in 3.5 for the r and b columns. vmstat 1 5 prints a line every 1 second, 5 times. Now read its memory and swap columns:
$ vmstat 1 5
procs -----------memory---------- ---swap-- -----io---- -system-- -------cpu-------
r b swpd free buff cache si so bi bo in cs us sy id wa st gu
1 0 12544 2143309 88020 2845980 0 0 31 12 128 201 1 1 98 0 0 0
1 0 12544 2143309 88020 2845980 0 0 0 4 197 331 1 1 98 0 0 0
1 0 12544 2143309 88020 2845980 0 0 0 4 204 342 1 1 98 0 0 0
Ignore the first line - it is the average since boot. Then:
rrunnable,bblocked in D state (the load average, split - 3.3).swpd,free,buff,cache- KiB in swap, free, buffers, cache.si/so- KiB/s swapped in / out during that interval. This is the one that matters. Sustained non-zerosiandsotogether is thrashing: what the processes need does not fit in RAM and the box spends its time moving pages between RAM and disk.wa- CPU idle waiting for I/O. Thrashing shows as highwaplussi/so.
Proving the cache is available
sync flush dirty pages first
echo 3 | sudo tee /proc/sys/vm/drop_caches
free -h
Dirty pages are cached pages that were changed in RAM and not yet written back to disk. sync writes them out, so they become clean and droppable. Writing a number into /proc/sys/vm/drop_caches tells the kernel to throw caches away.
buff/cache collapses, free shoots up - and available barely moves, because that memory was already counted as available. You have not gained anything; you have thrown away a cache that now has to be re-read from disk.
1 drops the page cache, 2 dentries and inodes, 3 both. Note the sudo tee (1.11): sudo echo 3 > /proc/sys/vm/drop_caches fails, because the redirection is done by your unprivileged shell. Never do this in production as a "fix". It is a demonstration.
Swap
Swap is disk used as overflow for RAM: the kernel moves anonymous pages nobody has used lately out to disk, and reads them back if they are needed. It does not add memory, it adds a slow tier.
swapon --show lists the swap areas:
$ swapon --show
NAME TYPE SIZE USED PRIO
/swap.img file 4G 12.3M -2
NAME the file or partition, TYPE file or partition, SIZE and USED, PRIO which swap area is used first. (/swap.img is the same file you met in the 2.36 incident.)
vm.swappiness (0-200 on current kernels, default 60) is how eagerly the kernel swaps anonymous pages out rather than dropping page cache. Lower means "keep processes in RAM, drop cache first". 10 is a common setting for a database.
sysctl reads and sets kernel settings:
$ sysctl vm.swappiness
vm.swappiness = 60
$ sudo sysctl -w vm.swappiness=10
vm.swappiness = 10
sysctl -w (write) lasts until reboot. Permanent: a file in /etc/sysctl.d/, e.g. echo 'vm.swappiness = 10' | sudo tee /etc/sysctl.d/60-swappiness.conf. sysctl is only a front end: vm.swappiness is the file /proc/sys/vm/swappiness.
A little swap in use is fine - often it is pages touched once at boot and never again. Swap churn (constant si/so in vmstat) is the pathology: the box grinds and everything crawls, which is usually worse than a fast OOM kill.
This is also why the kubelet in the 2.36 incident refused to start with swap on: it was written to assume memory limits are real RAM.
Later (Ch 15): Kubernetes required swap off on its machines for years, because its whole model of memory limits assumes RAM. Newer releases support swap, off by default and opt-in per machine.
Overcommit: why malloc succeeds and the OOM comes later
$ sysctl vm.overcommit_memory
vm.overcommit_memory = 0
$ grep -E 'CommitLimit|Committed_AS' /proc/meminfo
CommitLimit: 7227936 kB
Committed_AS: 2078922 kB
Linux hands out address space it does not have: that is overcommit. malloc(4 GB) on a 6 GB box with 2 GB free succeeds; pages are only found when the program touches them. Committed_AS is how much has been promised in total; CommitLimit is the ceiling used only in strict mode.
Mode 0 (heuristic, the default) refuses only absurd requests, 1 never refuses, 2 strict accounting against CommitLimit. With the default, "out of memory" is not an error your program gets back from malloc - it is the OOM killer arriving later, at whoever is biggest. That is the reason lesson 5.7 exists.
What you can now do
- Read
freeand say whether a box is really short of memory (available). - Tell harmless swap use from harmful swap traffic with
vmstat. - Read and change a kernel setting with
sysctl.