OnCallReady

Lesson 3.5 · Processes & Signals · 18 min read

Reading top and vmstat, line by line

In plain words

Imagine the dashboard of a car. The top row shows how long you have been driving and a trend of how hard the engine worked recently. Next come gauges: how much fuel, how much of the engine is busy, how much is idling at a red light, how much is taken by the passenger fiddling with the radio. Below is a list of every passenger and what they are consuming.

top is that dashboard. Line 1 is uptime and load, line 2 counts processes by state, line 3 splits CPU into us, sy, id, wa and st (stolen by the hypervisor), lines 4-5 are memory with avail Mem. vmstat 1 is the same dashboard as one line per second, where r is the queue for CPU and b is the number of passengers stuck waiting.

Someone pastes a screenshot of top: "is it the CPU?"

In an incident channel you will be handed top or vmstat output and asked what it means. Every line has a job; once you can read them, "CPU, memory, or stuck on I/O?" takes seconds instead of guesses.

What you need to know already: 3.1 (ps columns, STAT), 3.3 (load average, us/sy/id/wa, D state).

The five header lines

top - 20:00:04 up 2:17,  1 user,  load average: 0.09, 0.09, 0.05
Tasks:  46 total,   1 running,  45 sleeping,   0 stopped,   0 zombie
%Cpu(s):  0.7 us,  0.3 sy,  0.0 ni, 99.0 id,  0.0 wa,  0.0 hi,  0.0 si,  0.0 st
MiB Mem :   5925.1 total,   2083.5 free,    976.4 used,   2865.2 buff/cache
MiB Swap:   4096.0 total,   4083.7 free,     12.3 used.   4547.6 avail Mem

Line 1 is uptime: clock, time since boot, logged-in users, and the 1/5/15 minute load averages. Read the three numbers as a trend: 14.2, 11.6, 5.8 means "it got bad in the last few minutes and is still getting worse"; 0.9, 4.1, 6.3 means "it is recovering".

Line 2 counts processes by state. The two you scan for: stopped (someone hit Ctrl+Z or sent SIGSTOP and walked away) and zombie. "running" here means state R - on a CPU or waiting for one.

Line 3 splits CPU time, as a percentage of all cores:

us   user code            sy   kernel code           ni   user code at nice > 0
id   idle                 wa   idle WITH disk I/O outstanding
hi   hardware interrupts  si   software interrupts   st   stolen by the hypervisor

Lines 4-5 are memory (Chapter 5 goes deep). Swap is disk space the kernel uses as overflow when RAM is full. The number to read is avail Mem at the end of the swap line - how much memory programs could still get - not "free".

The blind spot in wa

wa is time a CPU sat idle while a task it ran was waiting for a disk (the kernel function is io_schedule()) - including waiting for writeback, which covers data headed for an NFS server. It is not "time spent in D state". A process blocked on an NFS request (wchan rpc_wait_bit_killable) is in D, adds 1 to the load, and contributes nothing to wa.

So: high load, high id, zero wa does not clear storage - it points at network filesystems. And high wa does not prove a local disk problem either. Count D-state processes and read their stacks instead of trusting wa.

The process columns

    PID USER      PR  NI    VIRT    RES    SHR S  %CPU  %MEM     TIME+ COMMAND
   1210 appuser   20   0 4200000 612000 171360 S   1.3  10.1   1:46.89 java
PR     the kernel's scheduling priority (20 = normal; rt = real-time)
NI     nice, -20 (greedy) .. 19 (polite). 0 by default.
VIRT   virtual size, KiB  = ps VSZ. Mostly meaningless.
RES    resident, KiB      = ps RSS. What is really in RAM.
SHR    the part of RES that is shared with other processes (libraries)
S      state: R S D Z T I
%CPU   of ONE core. 200% is possible on a 2-core box.
%MEM   RES as a share of total RAM
TIME+  CPU time used since start, minutes:seconds.hundredths

The "%CPU of one core" detail catches everyone: a Java process at 180% on a 2-core machine is using almost the whole box, while the %Cpu(s) line above would say about 90% us.

Keys worth knowing

P   sort by %CPU (default)       M   sort by memory
1   one line per CPU             c   show full command lines
k   kill: asks for PID and signal (default 15 = SIGTERM)
u   only one user's processes    o   filter, e.g. COMMAND=java
H   show threads                 q   quit

top in scripts and tickets

Interactive top is useless in a ticket or a log. Batch mode (-b) prints plain text instead of redrawing the screen; -n 1 = one snapshot, then exit:

$ top -b -n 1 | head -15

Paste that into an incident channel instead of a screenshot.

vmstat: the one-line summary

vmstat 1 5 prints one line of system-wide numbers every 1 second, 5 times:

$ vmstat 1 5
procs -----------memory---------- ---swap-- -----io---- -system-- -------cpu-------
 r  b   swpd   free   buff  cache   si   so    bi    bo   in   cs us sy id wa st gu
 1  0  12544 2137569  51200 2934000    0    0     3    11  104  180  1  0 99  0  0  0
 1  0  12544 2137569  51200 2934000    0    0     0     0   98  171  0  0 100 0  0  0

The quick triage: r high and id low = CPU. b high = something is stuck on I/O. so non-zero = RAM.

/proc/loadavg

$ cat /proc/loadavg
0.09 0.09 0.05 1/87 17395

The three averages, then running/total (counted in threads, not processes), then the most recently assigned PID. That last number climbing fast on an idle box means something is starting processes in a loop.

Normalise before you panic

Load is only meaningful relative to cores:

$ nproc
2
$ uptime
 20:00:03 up  2:17,  1 user,  load average: 3.90, 3.40, 2.10

3.9 on 2 cores is "twice as much runnable-or-blocked work as CPUs". 3.9 on 16 cores is idle. Always nproc first.

What you can now do

Why it helps

During an incident someone will paste a screenshot of top and ask "is it the CPU?". Reading every field quickly is what lets you answer correctly: st high on a virtual machine means the hypervisor is giving your CPU time to someone else, not that your code is slow; 180% on a 2-core box is the whole machine; wa zero with high load points at NFS.

top -b -n 1 and vmstat 1 5 are the text you put into tickets and incident channels instead of screenshots. vmstat's r, b and so columns give a CPU, I/O or memory answer in five seconds, which is a real advantage when everyone is guessing.

Commands in this lesson

top vmstat cat nproc uptime

FAQ

What is st (steal) and should I worry about it?

Steal is time your virtual CPU was ready to run but the hypervisor gave the physical CPU to something else: another VM on the same host, or on your laptop macOS itself. A few percent is normal on shared hosts. Sustained high steal means the host is oversubscribed. Your code cannot fix it; the fix is on the host side (more physical CPU, fewer VMs, a bigger VM size).

Why does the first line of vmstat look different from the rest?

The first line is averaged since boot, so it describes the machine's whole uptime, not now. Every following line covers the interval you asked for, such as one second with vmstat 1. Always ignore the first line when diagnosing. Several other sampling tools follow the same convention, so the habit pays off.

Can a process show more than 100% CPU?

Yes. In top's process list, %CPU is relative to one core by default ("Irix mode"), so a multi-threaded process using two cores fully shows 200%. The summary line %Cpu(s) is relative to all cores. Pressing Shift+I in top toggles Solaris mode, which divides by the number of cores. Keep this in mind when comparing numbers between tools.

Should I use htop instead of top?

For interactive use, htop is friendlier: per-core bars, mouse support, tree view, easier sorting and killing. Ubuntu Server ships it, but minimal installs and many other distributions do not, so you must know plain top anyway. For tickets and scripts, top -b -n 1 or ps and vmstat are better than any interactive tool.

What does a climbing last-PID number in /proc/loadavg mean?

The last field is the most recently assigned PID. On a quiet box it increases slowly. If it jumps by thousands per minute, something is starting processes constantly: a script in a tight loop, a check that runs a shell every second, a helper that crashes and is restarted. Each one is short-lived, so it may not show in top. Watching pstree -p or the journal around the parent usually finds it.

In an interview Junior

How do you tell quickly whether a box is CPU-bound, short on memory, or stuck on I/O?

vmstat 1 5, ignoring the first line (it is the average since boot):

Then top for the details: 1 for one line per core, P/M to sort, st for time a VM wanted a CPU but the hypervisor gave it to someone else. Remember %CPU is per core, so 180% on a 2-core box is almost all of it. Normalise load with nproc first, and for a ticket paste top -b -n 1 | head -15 rather than a screenshot.

Also asked: Explain the %Cpu(s) line in top. · What does st (steal) mean, and when would you see it? · Why does an NFS wait not show up in wa?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.