OnCallReady

Lesson 3.3 · Processes & Signals · 11 min read

Load average is not CPU usage

In plain words

Picture a post office with four counters. The queue count on the wall shows everyone standing at a counter or waiting in line. But this post office also counts people stuck at a counter waiting for a parcel to arrive from the warehouse, even though the clerk is just sitting there.

Load average is that number: processes running or wanting to run (R) plus processes stuck waiting in the kernel (D), averaged over 1, 5 and 15 minutes. Compare it with the number of counters, nproc. Load 15 with four counters and idle clerks means fifteen people waiting for parcels, like a dead NFS server. And nobody can pull a D person out of the line, not even kill -9, until their parcel arrives.

"Load is 14, add more CPUs!"

The most common alert on a Linux box is "load average is high". Half the time the CPUs are doing nothing, and adding CPUs changes nothing. This lesson is how to tell the two cases apart in under a minute.

What you need to know already: 3.1 (ps and the STAT letters R, S, D).

What the three numbers count

uptime prints the clock, how long the box has been up, how many users are logged in, and the load average:

$ uptime
 20:00:03 up  2:17,  1 user,  load average: 14.28, 11.67, 5.87

The load average is the average number of processes that were either R (running, or waiting for a CPU) or D (stuck waiting inside the kernel), over the last 1, 5 and 15 minutes.

That "or D" is the whole lesson. On Linux, unlike most other Unix systems, load includes processes blocked waiting for input/output (I/O: reading or writing a disk or a network file server). So:

A core is one CPU that can run one thing at a time; nproc prints how many this box has.

Rule of thumb: compare load with nproc, then immediately check whether the CPUs are actually busy. If they are not, count your D-state processes.

top, and the four keys worth knowing

top is ps that refreshes every few seconds and fills the screen. Keys:

  1   split the CPU line into one line per core
  M   sort by memory
  P   sort by CPU (the default)
  q   quit

Its CPU line splits all CPU time into percentages:

%Cpu(s):  3.1 us,  1.2 sy,  0.0 ni, 95.5 id,  0.2 wa,  0.0 hi,  0.0 si

High load + high us/sy means you really need more CPU. High load with us and sy near zero means processes are blocked, not computing.

Where the unused time shows up depends on what they wait for:

So never let wa decide for you: count the D-state processes (3.5 shows vmstat's b column, which counts them for you).

htop is a friendlier top (installed on Ubuntu Server): coloured bars per core, F6 to sort, F9 to send a signal. Same information.

You cannot kill a D

A process in D is inside a system call - a request a program makes to the kernel, like "read this file" - that did not ask to be interruptible. Signals are queued but not delivered until the call returns to the program, and it never returns while the disk or server it waits for never answers.

kill -9 returns success and does nothing at all. The process is not ignoring the signal; the signal has not been delivered yet.

The only ways out:

  1. The I/O completes (the file server comes back).
  2. Remove the cause - unmount the dead network directory with umount -f or umount -l (force / lazy), fix the network path.
  3. Reboot.

So finding a pile of D processes is not "what do I kill", it is "what are they all waiting for". /proc answers that: it is a folder the kernel fills with live information, one sub-folder per PID (3.14 tours it):

cat /proc/<pid>/wchan       the kernel function it is sleeping in
sudo cat /proc/<pid>/stack  the whole chain of kernel functions (root only)

(wchan = "wait channel". The kernel stack is the list of kernel functions that called each other to get there, innermost first.) How to read the names:

One refinement: many kernel waits are killable (the kernel calls them TASK_KILLABLE). They still show as D, but kill -9 is delivered. The tell is the wchan name: rpc_wait_bit_killable dies on kill -9; folio_wait_bit_common in writeback does not. So "D means unkillable" is the rule, and the wchan tells you whether you are looking at the exception.

What you can now do

Why it helps

"Load is 15, the box is dying" is one of the most common alerts, and half the time the CPU is idle. Knowing that Linux load includes D-state tasks turns that alert into the right question: what is everyone waiting for? The answer is usually a stuck NFS mount or a failing or overloaded disk, and restarting services or adding CPUs does nothing.

It also saves you from the kill -9 loop that wastes the first ten minutes of many incidents. Reading /proc/PID/wchan and stack tells you "NFS" or "block I/O" in seconds. This exact scenario is one of the Notion questions and a favourite in SRE interviews.

Commands in this lesson

nproc uptime

FAQ

Is a load of 4 bad?

It depends on the number of CPUs. Load 4 on a 4-core box means on average every core had one task: fully used, not overloaded. On a 32-core box it is nearly idle; on a 1-core box it means three tasks waiting at all times. Always divide by nproc first, then check whether the CPU is actually busy, since D-state tasks inflate load without using CPU.

Why is Linux load different from other Unixes?

Linux counts tasks in uninterruptible sleep (D) in addition to runnable ones; most other Unix systems count only runnable tasks. The decision dates from 1993, to reflect demand for resources beyond CPU, notably disk. It makes load a measure of "system demand" rather than CPU pressure, which is why you must look at CPU idle and D-state counts to interpret it.

Why does kill -9 return success on a D process but nothing happens?

kill only asks the kernel to queue the signal, which succeeds. Delivery happens when the process returns from the kernel to userspace. A D-state process is blocked inside a kernel call that is not interruptible, so it never gets there until the I/O completes. The signal stays pending; the process is not ignoring it. When the I/O finally finishes or fails, the pending SIGKILL is acted on.

How do I get processes out of D state without rebooting?

Fix what they wait for. If an NFS server is down, bring it back or restore the network path, and the calls complete. umount -f or a lazy umount -l of a dead NFS mount can let some calls fail. Mounts that use the soft option return errors after timeouts instead of hanging forever, at the risk of data corruption for writes. For a failed local disk, the device may need to be removed; often a reboot is the only way.

What do the three load numbers tell me?

They are exponentially damped averages over 1, 5 and 15 minutes. Read them as a trend: 14, 11, 5 means it rose recently and is still climbing; 1, 5, 9 means it is recovering. A single high 1-minute value can be a short burst like a cron job; a high 15-minute value means sustained load. They lag reality, so confirm with vmstat 1.

In an interview Junior

Load average is 15 on a 4-core machine, but the CPU is 90% idle. What is going on?

On Linux the load average counts processes that are R (running or waiting for a CPU) or D (stuck waiting inside the kernel, usually for I/O). With the CPU idle, it is not a CPU problem: about fifteen processes are blocked on something - an overloaded disk or a dead network file server (NFS). Adding CPUs changes nothing.

How I'd check: nproc and uptime to compare load with cores; top to confirm us/sy are near zero; then count the D processes - ps -eo pid,stat,comm or vmstat's b column. Then ask what they wait for: cat /proc/PID/wchan and sudo cat /proc/PID/stack; names like nfs_* and rpc_* mean NFS. Don't trust wa to decide - an NFS wait shows as plain idle. kill -9 won't clear them; the way out is the I/O completing, unmounting the dead share, or a reboot.

Also asked: A process will not die even with kill -9. What state is it in, and what does that tell you? · What do the three numbers of the load average mean? · What is the difference between load average and CPU utilisation?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.