Someone pastes a screenshot of top: "is it the CPU?"
In an incident channel you will be handed top or vmstat output and asked what it means. Every line has a job; once you can read them, "CPU, memory, or stuck on I/O?" takes seconds instead of guesses.
What you need to know already: 3.1 (ps columns, STAT), 3.3 (load average, us/sy/id/wa, D state).
The five header lines
top - 20:00:04 up 2:17, 1 user, load average: 0.09, 0.09, 0.05
Tasks: 46 total, 1 running, 45 sleeping, 0 stopped, 0 zombie
%Cpu(s): 0.7 us, 0.3 sy, 0.0 ni, 99.0 id, 0.0 wa, 0.0 hi, 0.0 si, 0.0 st
MiB Mem : 5925.1 total, 2083.5 free, 976.4 used, 2865.2 buff/cache
MiB Swap: 4096.0 total, 4083.7 free, 12.3 used. 4547.6 avail Mem
Line 1 is uptime: clock, time since boot, logged-in users, and the 1/5/15 minute load averages. Read the three numbers as a trend: 14.2, 11.6, 5.8 means "it got bad in the last few minutes and is still getting worse"; 0.9, 4.1, 6.3 means "it is recovering".
Line 2 counts processes by state. The two you scan for: stopped (someone hit Ctrl+Z or sent SIGSTOP and walked away) and zombie. "running" here means state R - on a CPU or waiting for one.
Line 3 splits CPU time, as a percentage of all cores:
us user code sy kernel code ni user code at nice > 0
id idle wa idle WITH disk I/O outstanding
hi hardware interrupts si software interrupts st stolen by the hypervisor
- An interrupt is a device (network card, disk) tapping the CPU on the shoulder: "data arrived, deal with it".
hi/siis time spent doing that. - nice is a process's politeness, from -20 (greedy) to 19 (polite).
st(steal): this box is a virtual machine; the hypervisor is the program on the real hardware that runs the VMs (UTM on your Mac, 1.1).stis time your VM wanted a CPU but the hypervisor gave the real CPU to someone else - another VM, or macOS itself. It does not show up inus.
Lines 4-5 are memory (Chapter 5 goes deep). Swap is disk space the kernel uses as overflow when RAM is full. The number to read is avail Mem at the end of the swap line - how much memory programs could still get - not "free".
The blind spot in wa
wa is time a CPU sat idle while a task it ran was waiting for a disk (the kernel function is io_schedule()) - including waiting for writeback, which covers data headed for an NFS server. It is not "time spent in D state". A process blocked on an NFS request (wchan rpc_wait_bit_killable) is in D, adds 1 to the load, and contributes nothing to wa.
So: high load, high id, zero wa does not clear storage - it points at network filesystems. And high wa does not prove a local disk problem either. Count D-state processes and read their stacks instead of trusting wa.
The process columns
PID USER PR NI VIRT RES SHR S %CPU %MEM TIME+ COMMAND
1210 appuser 20 0 4200000 612000 171360 S 1.3 10.1 1:46.89 java
PR the kernel's scheduling priority (20 = normal; rt = real-time)
NI nice, -20 (greedy) .. 19 (polite). 0 by default.
VIRT virtual size, KiB = ps VSZ. Mostly meaningless.
RES resident, KiB = ps RSS. What is really in RAM.
SHR the part of RES that is shared with other processes (libraries)
S state: R S D Z T I
%CPU of ONE core. 200% is possible on a 2-core box.
%MEM RES as a share of total RAM
TIME+ CPU time used since start, minutes:seconds.hundredths
The "%CPU of one core" detail catches everyone: a Java process at 180% on a 2-core machine is using almost the whole box, while the %Cpu(s) line above would say about 90% us.
Keys worth knowing
P sort by %CPU (default) M sort by memory
1 one line per CPU c show full command lines
k kill: asks for PID and signal (default 15 = SIGTERM)
u only one user's processes o filter, e.g. COMMAND=java
H show threads q quit
top in scripts and tickets
Interactive top is useless in a ticket or a log. Batch mode (-b) prints plain text instead of redrawing the screen; -n 1 = one snapshot, then exit:
$ top -b -n 1 | head -15
Paste that into an incident channel instead of a screenshot.
vmstat: the one-line summary
vmstat 1 5 prints one line of system-wide numbers every 1 second, 5 times:
$ vmstat 1 5
procs -----------memory---------- ---swap-- -----io---- -system-- -------cpu-------
r b swpd free buff cache si so bi bo in cs us sy id wa st gu
1 0 12544 2137569 51200 2934000 0 0 3 11 104 180 1 0 99 0 0 0
1 0 12544 2137569 51200 2934000 0 0 0 0 98 171 0 0 100 0 0 0
- The first line is the average since boot. Ignore it; read the rest.
r- runnable processes (R). Sustained r > number of cores = CPU contention.b- processes blocked in D state. This is the column that answers "is the load average I/O?" withoutpsgymnastics.swpd,free,buff,cache- memory in KiB (Chapter 5).si/so- swap in/out per second. Sustainedso> 0 = RAM is short.bi/bo- blocks read from / written to disk per second.in/cs- interrupts and context switches (the CPU switching from one process to another) per second.- the CPU columns are the same as top's;
guis time spent running guest VMs.
The quick triage: r high and id low = CPU. b high = something is stuck on I/O. so non-zero = RAM.
/proc/loadavg
$ cat /proc/loadavg
0.09 0.09 0.05 1/87 17395
The three averages, then running/total (counted in threads, not processes), then the most recently assigned PID. That last number climbing fast on an idle box means something is starting processes in a loop.
Normalise before you panic
Load is only meaningful relative to cores:
$ nproc
2
$ uptime
20:00:03 up 2:17, 1 user, load average: 3.90, 3.40, 2.10
3.9 on 2 cores is "twice as much runnable-or-blocked work as CPUs". 3.9 on 16 cores is idle. Always nproc first.
What you can now do
- Read every line of top's header, including
standavail Mem. - Use
vmstat 1 5to answer "CPU, I/O or RAM?" fromr,bandso. - Capture a text snapshot for a ticket with
top -b -n 1.