"Load is 14, add more CPUs!"
The most common alert on a Linux box is "load average is high". Half the time the CPUs are doing nothing, and adding CPUs changes nothing. This lesson is how to tell the two cases apart in under a minute.
What you need to know already: 3.1 (ps and the STAT letters R, S, D).
What the three numbers count
uptime prints the clock, how long the box has been up, how many users are logged in, and the load average:
$ uptime
20:00:03 up 2:17, 1 user, load average: 14.28, 11.67, 5.87
The load average is the average number of processes that were either R (running, or waiting for a CPU) or D (stuck waiting inside the kernel), over the last 1, 5 and 15 minutes.
That "or D" is the whole lesson. On Linux, unlike most other Unix systems, load includes processes blocked waiting for input/output (I/O: reading or writing a disk or a network file server). So:
- Load 15 on 4 cores with the CPU 90% idle is not a CPU problem. It is fifteen processes stuck waiting on something - a dead network file server, an overloaded disk.
- Load 4 on 4 cores with 100% CPU is a perfectly healthy, fully used box.
A core is one CPU that can run one thing at a time; nproc prints how many this box has.
Rule of thumb: compare load with nproc, then immediately check whether the CPUs are actually busy. If they are not, count your D-state processes.
top, and the four keys worth knowing
top is ps that refreshes every few seconds and fills the screen. Keys:
1 split the CPU line into one line per core
M sort by memory
P sort by CPU (the default)
q quit
Its CPU line splits all CPU time into percentages:
%Cpu(s): 3.1 us, 1.2 sy, 0.0 ni, 95.5 id, 0.2 wa, 0.0 hi, 0.0 si
ususer: running programs' own codesysystem: running kernel code on their behalfninice: user code of low-priority processesididle: nothing to dowaiowait: idle, but some process it ran is waiting for disk I/Ohi,si: handling hardware and software interrupts (3.5)
High load + high us/sy means you really need more CPU. High load with us and sy near zero means processes are blocked, not computing.
Where the unused time shows up depends on what they wait for:
- Waiting for a local disk, or for writeback - the kernel writing data a program already "saved" into RAM out to the real disk or server - counts as
wa. - Waiting for an answer from NFS (Network File System: a directory that really lives on another server, reached over the network) counts as plain
id, while still adding 1 to the load.
So never let wa decide for you: count the D-state processes (3.5 shows vmstat's b column, which counts them for you).
htop is a friendlier top (installed on Ubuntu Server): coloured bars per core, F6 to sort, F9 to send a signal. Same information.
You cannot kill a D
A process in D is inside a system call - a request a program makes to the kernel, like "read this file" - that did not ask to be interruptible. Signals are queued but not delivered until the call returns to the program, and it never returns while the disk or server it waits for never answers.
kill -9 returns success and does nothing at all. The process is not ignoring the signal; the signal has not been delivered yet.
The only ways out:
- The I/O completes (the file server comes back).
- Remove the cause - unmount the dead network directory with
umount -forumount -l(force / lazy), fix the network path. - Reboot.
So finding a pile of D processes is not "what do I kill", it is "what are they all waiting for". /proc answers that: it is a folder the kernel fills with live information, one sub-folder per PID (3.14 tours it):
cat /proc/<pid>/wchan the kernel function it is sleeping in
sudo cat /proc/<pid>/stack the whole chain of kernel functions (root only)
(wchan = "wait channel". The kernel stack is the list of kernel functions that called each other to get there, innermost first.) How to read the names:
nfs_*,rpc_*,[sunrpc]- NFS (RPC = remote procedure call, how NFS asks the server for things).io_schedule,blk_mq_*,ext4_*- a local disk.folio_wait_bit_common/folio_wait_writeback- waiting for writeback, to whatever filesystem the rest of the stack names.
One refinement: many kernel waits are killable (the kernel calls them TASK_KILLABLE). They still show as D, but kill -9 is delivered. The tell is the wchan name: rpc_wait_bit_killable dies on kill -9; folio_wait_bit_common in writeback does not. So "D means unkillable" is the rule, and the wchan tells you whether you are looking at the exception.
What you can now do
- Read
uptime's three numbers againstnproc. - Tell "out of CPU" (high
us/sy) from "stuck on I/O" (idle CPU, many D). - Find what a D process waits for from
/proc/<pid>/wchanandstack.