OnCallReady

Lesson 0.5 · SRE Fundamentals · 21 min read

Saturation and the USE method

In plain words

Think of a single checkout lane at a supermarket. You can ask three questions about it. How busy is the cashier? (Utilisation.) How many people are standing in line waiting? (Saturation.) How often does the till jam? (Errors.) A cashier who is busy all the time with nobody waiting is fine. The moment people are queuing, everyone waits longer, even if the cashier works no harder.

The USE method, from the performance engineer Brendan Gregg, asks those three questions for every resource: the CPU (how busy, and how much work is waiting for it), memory (how much is left), disk (how full, and how fast it fills), file descriptors (the handles a program needs for every open file or connection: open versus limit), and the database connection pool (in use versus max, and requests waiting). Little's law, in-flight = rate x time, tells you when a pool must run out.

Why this lesson

During checkout's outage the processor graph looked calm - lower than usual, even. The team's only alert watched the processor, so it stayed silent while a third of customers got errors. The resource that had run out was a different one, and nobody was looking at it. This lesson gives you a method for finding the resource that runs out first, every time.

What you need to know already: 0.1 The four golden signals (saturation, connection pool), 0.3 Golden signals from an access log (process, PID, /proc, file descriptors).

Utilisation, saturation, errors

Brendan Gregg's USE method: for every resource (anything a service needs a share of - the processor, memory, disk, a pool of database connections), check three things.

utilisation   how busy it is: % of time busy, or used / capacity
saturation    how much work is WAITING for it: queue length, requests pending
errors        how often it fails: error counts, "too many open files", timeouts

The golden signals tell you that a service hurts. USE, walked resource by resource, tells you which resource is the reason. Saturation is the column that matters most, because a resource can be 100% utilised and fine (a processor doing a big calculation that nobody is waiting for), but the moment work is queueing for it, every request pays the wait.

Walking it on this box

Each command below is a standard Linux tool that prints a snapshot of one resource. You will meet them all properly later; here, just learn to read them.

CPU. nproc prints how many CPUs (processor cores - each can run one program at a time) the box has. uptime prints how long it has been running and the load average. vmstat 1 2 ("virtual memory statistics") prints a line of counters every 1 second, 2 times.

$ nproc
2
$ uptime
 20:00:03 up  2:17,  1 user,  load average: 0.09, 0.09, 0.05
$ vmstat 1 2
procs -----------memory---------- ---swap-- -----io---- -system-- -------cpu-------
 r  b   swpd   free   buff  cache   si   so    bi    bo   in   cs us sy id wa st gu
 1  0  12544 2143309  88020 2845980    0    0    31    12  128  201  1  1 98  0  0  0
 1  0  12544 2143309  88020 2845980    0    0     0     4  197  331  1  1 98  0  0  0

The columns that matter here: r = tasks ready to run and waiting for a CPU; us + sy = % of CPU time spent in programs and in the system (utilisation); id = % idle; wa = % waiting for the disk. The load average (the three numbers: over the last 1, 5 and 15 minutes) is roughly the average number of tasks that wanted a CPU, or were stuck waiting for the disk.

Two CPUs, r = 1, 98% idle: not saturated. r consistently above the CPU count means work is waiting for a CPU. A high load with idle CPUs means tasks are stuck waiting for something else - usually the disk or the network. The first vmstat line is an average since the box started; read from the second line on.

Memory. free -m shows memory in megabytes (-m):

$ free -m
               total        used        free      shared  buff/cache   available
Mem:            5925         966        2093           4        2865        4557
Swap:           4095          12        4083

total is all the memory; used is taken by programs; buff/cache is memory the system borrows to keep recently read files handy (it gives it back when programs need it); available is what programs could still get - the column that matters. Swap is disk space used as overflow memory; it is far slower, so heavy swapping (si/so above zero in vmstat) is memory saturation. At the very end, the system kills a program to free memory.

Disk space. df -h ("disk free", -h = human-readable sizes like 19G) for the two disks mounted at / and /data:

$ df -h / /data
Filesystem                         Size  Used Avail Use% Mounted on
/dev/mapper/ubuntu--vg-ubuntu--lv   19G  7.5G  9.7G  44% /
/dev/vdb1                           50G   42G  6.2G  87% /data

Read it by the last column: the disk you see at / is 19 GB, 44% used; the one at /data is 87% full. (The first column is the system's name for each disk

alert is a prediction ("full in 4 hours") rather than a threshold. 87% on /data means nothing on its own: is it growing 1 GB a day (a week of headroom) or 1 GB an hour?

Disk speed. Utilisation is how busy the disk is; saturation is how many reads and writes are queued for it (iostat -x, not installed here, shows both) and wa in vmstat.

File descriptors. Utilisation is open / limit; the error is EMFILE (Too many open files). There is no queue: at 100% the next attempt to open a file or accept a connection fails.

$ grep 'open files' /proc/$(pgrep -f orders.jar)/limits
Max open files            1024                 1024                 files

(grep prints the matching line of the orders process's limits file; the two numbers are the limit in force and the ceiling it could be raised to.)

The resources nobody graphs

For a web service, the resources that run out first are rarely the CPU:

ResourceUtilisationSaturationErrors
Database connection poolconnections in use / maximumrequests pending (waiting for a connection)"connection not available" timeouts
Request workersbusy / maximum (often 200)connections queued waiting for a workerconnection refused
Memoryused / limittime the program spends cleaning up memory instead of workingout-of-memory crash
CPU, when a limit is set on the serviceusage / limittime the service was held back by the limit-
File descriptorsopen / limit-EMFILE

Checkout's outage (#4471 - incidents get a number so everyone can refer to the same one) is the textbook case. The lab has a per-minute export of checkout's pool. The command prints, per minute: connections in use / maximum, requests pending, and CPU in use (in cores: 0.5 = half of one CPU):

$ cd ~/oncall-lab/labs/0-sre
$ awk -F'\t' '!/^#/ {printf "%s %3d/%3d pending=%s cpu=%s\n", substr($1,12,5), $2, $3, $4, $7}' data/checkout-jvm.tsv | sed -n '26,33p'
18:27  12/120 pending=0 cpu=0.48
18:28  16/120 pending=0 cpu=0.49
18:29  24/ 90 pending=3 cpu=0.58
18:30  26/ 60 pending=3 cpu=0.53
18:31  30/ 30 pending=6 cpu=0.43
18:32  30/ 30 pending=15 cpu=0.41
18:33  30/ 30 pending=29 cpu=0.40
18:34  30/ 30 pending=20 cpu=0.40

(cd moves you into a directory. !/^#/ skips lines starting with #, the file's comment header. substr($1,12,5) takes 5 characters of column 1 starting at character 12 - the HH:MM.)

Your clock times differ; the shape does not. Checkout runs as 3 instances. As each one picks up the new release 2.8.0, the pool maximum across all three falls from 120 to 30 (40 each, then 10 each). Requests start waiting (pending) at 18:29 - three minutes before the first 5xx. CPU goes down: requests waiting for a connection do not use the processor; it drifts from about 0.5 cores to 0.4 and below. A CPU alert watches the one resource that got quieter.

Little's law: why a pool of 10 runs out

For anything where work arrives, waits and leaves (a queue), Little's law says:

L = lambda x W
in-flight items = arrival rate x time each item spends in the system

Applied to a connection pool: connections in use = requests per second that need a connection x how long each holds it.

$ awk 'BEGIN {for (rps=50; rps<=400; rps*=2) printf "%3d req/s x 40 ms = %4.1f connections busy on average\n", rps, rps*0.040}'
 50 req/s x 40 ms =  2.0 connections busy on average
100 req/s x 40 ms =  4.0 connections busy on average
200 req/s x 40 ms =  8.0 connections busy on average
400 req/s x 40 ms = 16.0 connections busy on average

At 200 requests per second a pool of 10 is 80% utilised on average - and traffic comes in bursts, so it is regularly full. At 400 req/s it cannot keep up at any moment, the queue grows, and every queued request waits until the pool's timeout gives up on it. Little's law also runs the other way: if database queries slow down (W doubles), the same traffic needs twice the connections. (A dependency is another service yours needs in order to work - here, the database.) That is why a slow database shows up as a pool problem in the service, and why "make the pool bigger" can make things worse - more queries at once on a database that is already slow make each one slower.

Saturation is a leading indicator

Latency, errors and traffic describe what already happened. Saturation describes what is about to happen, because most systems degrade before 100%:

So alert on saturation before the cliff: pending > 0 for 2 minutes, r above the CPU count for 10 minutes, disk predicted full in 4 hours. And send it as a ticket (a task in the team's to-do system, handled in working hours) or show it as supporting evidence, not as a page (an alert that wakes someone up) on its own: saturation that does not hurt users yet is a job for working hours - unless the prediction says it will hurt them tonight.

The averaging trap, again

Added up across the three instances, the pool looked fine while only one of them was on 2.8.0: 24/90 = 27% in use. But that one instance was at 10/10 with requests waiting; the other two were nearly idle. Saturation is per instance. nginx does not move a waiting request to an idle instance - each request is stuck on the one it landed on. Look at the maximum across instances, not the sum or the average.

In short

USE           every resource: utilisation, saturation (waiting), errors
leading       saturation warns before latency and errors move
web service   pool pending, busy workers, memory, fds - not just CPU
Little        in-flight = rate x time; slow dependencies eat pools
per instance  the maximum across instances, never the sum

What you can now do:

Why it helps

When a service is slow and the golden signals say "it hurts", USE is how you find which resource is the reason. Situations: checkout at 0.4 cores and failing, with pending=29 on the pool, while a CPU alert watches the one thing that got quieter. Someone proposes making the pool bigger; Little's law says a slow database needs more connections for the same traffic, and more concurrent queries can make it slower. A pool graph summed over all instances at 27% hides one instance at 10/10; saturation is per instance. And the same questions work on any resource you meet later, from disks to memory. This is the diagnostic method interviewers want to hear when they ask "the service is slow, what do you do?"

Commands in this lesson

nproc uptime vmstat free df grep cd awk

FAQ

What is the difference between utilisation and saturation?

Utilisation is how busy a resource is: percent of time busy, or used over capacity. Saturation is how much work is waiting for it: run queue, pending threads, queue length. A resource can be at 100% utilisation and fine, like a CPU doing batch work with nothing waiting. Once work queues, every request pays the wait, so saturation is the column that matters most.

What does CPU saturation look like?

More work wants to run than there are processor cores to run it, so some of it waits in a queue. A machine with 2 cores and 2 busy programs is fully utilised but not saturated; with 6 programs wanting to run, 4 wait every moment and everything gets slower. The number to watch is the length of that queue compared with the number of cores, not the busy percentage.

What does Little's law tell me about connection pools?

In-flight connections equal requests per second needing the database times how long each holds a connection. At 200 req/s and 40 ms, 8 connections are busy on average, so a pool of 10 is 80% utilised and regularly full under bursts. If queries slow down, the same traffic needs proportionally more connections, which is why a slow database appears as pool exhaustion in the application.

Why look at the maximum across instances, not the sum?

Saturation is per instance. A load balancer (the component that spreads requests over the instances) does not move a waiting request to an idle instance; it waits on the one where it landed. Summed across three instances the pool looked 27% used, while one instance was at 10/10 with requests pending. Take the highest instance for saturation signals.

Should saturation alerts page?

Usually not on their own. Saturation that doesn't hurt users yet is working-hours work: alert as a ticket or a supporting signal, with thresholds below the cliff, like pending greater than zero for 2 minutes or a disk predicted full in 4 hours. The exception is a prediction that says users will be hurt tonight, like a disk filling before morning.

In an interview Junior

What is the USE method?

A checklist for finding which resource is the problem: for every resource, check Utilisation, Saturation and Errors.

On a box I walk it resource by resource: CPU with vmstat 1 2 (r = waiting for a CPU), memory with free -m (the available column, and swap in use), disk with df -h, open files against their limit, and then the application's own resources, like a database pool's connections in use out of its maximum. The golden signals say a service hurts; USE says which resource is the reason. In checkout's outage the pool was at 30/30 with requests waiting while the CPU was nearly idle.

Also asked: What is the difference between utilisation and saturation? · A service is slow but its CPU is at 15%. What else could be saturated? · What does Little's law tell you about a connection pool?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.