OnCallReady

Lesson 4.19 · Filesystem, Permissions, Disk · 14 min read

du, df and finding the space

In plain words

Imagine two ways of checking how full a warehouse is. One is asking the manager at the door, who keeps a running total of every shelf rented, including shelves rented to people who threw away their key but never checked out. The other is walking the aisles yourself and adding up the boxes you can see and read labels on.

df is asking the manager: fast and complete. du is walking the aisles: slow, and it misses boxes behind doors you cannot open (no sudo) or boxes with no label (deleted but still open). When the two numbers disagree, that gap is exactly what you are looking for. du -xh --max-depth=2 | sort -h | tail is the walk, organised so the biggest aisles come last.

Why you need two tools for one question

"The disk is full - what is using it?" is one of the most common pages there is. Two commands answer it from opposite ends, and when they disagree, the disagreement is the clue.

What you need to know already: 4.18 (mounts, df columns), 4.13 (names vs inodes), 4.15 (find tests), 1.7 (pipes, 2>/dev/null).

They answer different questions

df ("disk free") asks the filesystem how many blocks are allocated. Fast, and it counts everything - including data whose name has been removed but which is still open.

du ("disk usage") walks the directory tree and adds up the files it can see. Slow, and it misses anything it cannot read, plus anything with no name left.

When they disagree, that gap is the finding.

The one-liner worth memorising

sudo du -xh / --max-depth=2 2>/dev/null | sort -h | tail -20

Then descend: run it again rooted at whatever came top.

$ sudo du -xh /data --max-depth=1 2>/dev/null | sort -h
4.0K	/data/backups
4.0K	/data/lost+found
41G	/data
41G	/data/app

Two columns: size, path. The line for the start directory itself is the total. Here one subdirectory is the whole of it, so the next command is du -xh /data/app --max-depth=1 - or, when you suspect a single file, find. (lost+found is a directory every ext4 filesystem has, used by the repair tool.)

What goes wrong without the flags

$ du -sh /data
du: cannot read directory '/data/backups': Permission denied
du: cannot read directory '/data/lost+found': Permission denied
41G	/data

(-s = one summary line per argument.) Without sudo, du still prints a total - of what it could read. Those error lines are telling you the number is a lower bound. And the difference between sort -n and sort -h:

$ printf '900M\n2.0G\n12K\n' | sort -n       $ printf '900M\n2.0G\n12K\n' | sort -h
2.0G                                        12K
12K                                         900M
900M                                        2.0G

sort -n reads the leading number and ignores the unit. You would "find" the 900M directory and miss the 2G one.

Apparent size versus disk usage

du reports blocks allocated, at least one 4 KiB block per file. ls -l reports the apparent size: the length in bytes. They disagree in both directions:

du --apparent-size switches du to the ls view. When the question is "why is the disk full", blocks are the truth.

find, for the big single files

sudo find /data -size +1G -type f
sudo find /data -size +1G -type f -printf '%s %p\n' | sort -rn
sudo find /var/log -name '*.gz' -mtime +30 -delete

-size +1G is "greater than 1 GiB". -mtime +30 is "modified more than 30 days ago". -printf '%s %p\n' prints the size in bytes before the path, so sort -rn (numeric, reversed: biggest first) can order them - bytes, not human units, because sort -n needs plain numbers.

Always run a -delete as a print first. find ... -print, read it, then change -print to -delete. There is no undo.

Why find beats rm at scale

# in the full-cache mission below, with 12,000 files in sessions/
sudo rm /srv/cache/sessions/*
bash: /usr/bin/rm: Argument list too long

The shell expands the glob * into every filename before rm runs, and the kernel limits how big a program's argument list may be (about 2MB; getconf ARG_MAX prints the limit). Twelve thousand filenames exceed it. find never builds that list - it walks and acts:

sudo find /srv/cache/sessions -type f -mtime +7 -delete

which also lets you be selective about which files, instead of all of them. rm -r /srv/cache/sessions would also work (no glob), but deletes the directory too, and the application may not recreate it.

What you can now do

Why it helps

"The disk is full" is one of the most frequent pages in operations, and the du -x | sort -h one-liner, repeated downwards, finds the culprit in minutes: a runaway log, an old backup tarball, a forgotten heap dump, a core dump. It is the literal answer to the Notion question about the top 10 largest directories excluding other mounts.

The details matter under pressure: forgetting -x gives nonsense numbers from /proc and other mounts, sort -n misorders human sizes, running without sudo gives a lower bound, and rm * in a directory with 12 thousand files fails with "Argument list too long" at the worst moment. Knowing find -delete and apparent size versus blocks also informs capacity planning for many small files.

Commands in this lesson

du printf

FAQ

Why use sort -h and not sort -n with du -h?

sort -n reads only the leading number and ignores the unit, so 900M sorts above 2.0G because 900 is larger than 2. sort -h understands human-readable suffixes (K, M, G, T) and sorts by the real size. Pair du -h with sort -h, or use du -k or bytes with sort -n. Both put the largest entries last, hence tail.

What does -x do in du and why is it essential?

-x (--one-file-system) makes du skip directories on other filesystems than the start point. Without it, du / descends into /proc, /sys, /run, other disks like /data, and network mounts, producing errors, huge times and numbers that do not correspond to the full filesystem you are investigating. Always use it when chasing a specific full filesystem.

Why is du's number different from ls -l sizes?

ls -l shows the apparent size, the length of the file in bytes. du shows allocated blocks, rounded to the block size (usually 4 KiB). Many tiny files use far more disk than their apparent sizes add up to, while sparse files like VM images or database preallocations use less. du --apparent-size shows the ls view; for "why is the disk full", allocated blocks are the truth.

Why does rm * fail with Argument list too long?

The shell expands * into every matching filename and passes them as arguments to rm. The kernel limits the total size of arguments and environment for one exec (ARG_MAX, a few megabytes on Linux, see getconf ARG_MAX). Thousands of long names exceed it. find dir -type f -delete or find ... -exec rm {} + avoids building the list, and lets you filter which files to delete.

Is there a faster way than repeating du by hand?

ncdu -x / scans once and lets you browse interactively, sorted by size, and delete from inside it; it is worth installing on servers you manage. du -xh --max-depth=3 / | sort -h | tail -30 gives a deeper single pass. For logs specifically, journalctl --disk-usage, and journalctl --vacuum-size= to shrink the journal safely.

In an interview Junior

How would you list the 10 largest directories on a server, without crossing into other filesystems?

sudo du -xh / --max-depth=2 2>/dev/null | sort -h | tail -10

Then descend into the top entry with the same command, or use sudo find /data -size +1G -type f for single big files. If du's total is far below what df says is used, the space is held by deleted-but-open files (lsof +L1) or hidden under a mount point.

Also asked: df says 42 GB used but du finds 8 KB. What is happening? · Why does rm /srv/cache/sessions/* fail with "Argument list too long", and what do you use instead? · What is the difference between apparent size and disk usage?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.