Why you need two tools for one question
"The disk is full - what is using it?" is one of the most common pages there is. Two commands answer it from opposite ends, and when they disagree, the disagreement is the clue.
What you need to know already: 4.18 (mounts, df columns), 4.13 (names vs inodes), 4.15 (find tests), 1.7 (pipes, 2>/dev/null).
They answer different questions
df ("disk free") asks the filesystem how many blocks are allocated. Fast, and it counts everything - including data whose name has been removed but which is still open.
du ("disk usage") walks the directory tree and adds up the files it can see. Slow, and it misses anything it cannot read, plus anything with no name left.
When they disagree, that gap is the finding.
The one-liner worth memorising
sudo du -xh / --max-depth=2 2>/dev/null | sort -h | tail -20
-xstay on one filesystem. Without it you walk into /proc, /sys and every other mount and the numbers are meaningless.-hhuman sizes (K, M, G), andsort -hsorts them correctly (2K < 1M < 1G) - a plainsort -nwould not.--max-depth=2print totals only two directory levels down, so the output stays readable.2>/dev/nullhides the permission-denied noise. Usesudoso there is less of it to hide.tailbecausesortputs the big ones last.
Then descend: run it again rooted at whatever came top.
$ sudo du -xh /data --max-depth=1 2>/dev/null | sort -h
4.0K /data/backups
4.0K /data/lost+found
41G /data
41G /data/app
Two columns: size, path. The line for the start directory itself is the total. Here one subdirectory is the whole of it, so the next command is du -xh /data/app --max-depth=1 - or, when you suspect a single file, find. (lost+found is a directory every ext4 filesystem has, used by the repair tool.)
What goes wrong without the flags
$ du -sh /data
du: cannot read directory '/data/backups': Permission denied
du: cannot read directory '/data/lost+found': Permission denied
41G /data
(-s = one summary line per argument.) Without sudo, du still prints a total - of what it could read. Those error lines are telling you the number is a lower bound. And the difference between sort -n and sort -h:
$ printf '900M\n2.0G\n12K\n' | sort -n $ printf '900M\n2.0G\n12K\n' | sort -h
2.0G 12K
12K 900M
900M 2.0G
sort -n reads the leading number and ignores the unit. You would "find" the 900M directory and miss the 2G one.
Apparent size versus disk usage
du reports blocks allocated, at least one 4 KiB block per file. ls -l reports the apparent size: the length in bytes. They disagree in both directions:
- a million 100-byte files:
lssays 100 MB total, du says ~4 GB (one block each); - a sparse file - one with holes that were never written, like a VM disk image or
truncate -s 10G-lssays 10G, du says what was actually written.
du --apparent-size switches du to the ls view. When the question is "why is the disk full", blocks are the truth.
find, for the big single files
sudo find /data -size +1G -type f
sudo find /data -size +1G -type f -printf '%s %p\n' | sort -rn
sudo find /var/log -name '*.gz' -mtime +30 -delete
-size +1G is "greater than 1 GiB". -mtime +30 is "modified more than 30 days ago". -printf '%s %p\n' prints the size in bytes before the path, so sort -rn (numeric, reversed: biggest first) can order them - bytes, not human units, because sort -n needs plain numbers.
Always run a -delete as a print first. find ... -print, read it, then change -print to -delete. There is no undo.
Why find beats rm at scale
# in the full-cache mission below, with 12,000 files in sessions/
sudo rm /srv/cache/sessions/*
bash: /usr/bin/rm: Argument list too long
The shell expands the glob * into every filename before rm runs, and the kernel limits how big a program's argument list may be (about 2MB; getconf ARG_MAX prints the limit). Twelve thousand filenames exceed it. find never builds that list - it walks and acts:
sudo find /srv/cache/sessions -type f -mtime +7 -delete
which also lets you be selective about which files, instead of all of them. rm -r /srv/cache/sessions would also work (no glob), but deletes the directory too, and the application may not recreate it.
What you can now do
- Find the biggest directories on one filesystem with
du -x | sort -h. - Find single huge files with
find -size. - Explain why
duanddfcan disagree, and whyrm dir/*can fail.