Why "the disk has space" does not end the conversation
A write fails with No space left on device, and df -h says the disk is half empty. Unless you know the three ways that can happen - and which to check first - you lose an hour, or someone adds disk that changes nothing.
What you need to know already: 4.18 (df -h, df -i, tune2fs -l), 4.19 (du vs df), 4.13 (link count, deleted-but-open files), 3.14 (/proc/<pid>/fd).
Three causes, and the order to check them
All three produce the same error text, No space left on device. Its short name is ENOSPC - every kernel error has a code like this (called an errno), and the text is just its translation.
1. Inodes exhausted. df -i
Every file costs one inode, however small. A filesystem is created with a fixed number of them, and millions of tiny files - web session files, a cache directory, a mail queue - use up the inodes long before the blocks. This is inode exhaustion:
$ df -h /srv/cache $ df -i /srv/cache
Size Used Avail Use% Inodes IUsed IFree IUse%
2.0G 802M 1.2G 42% 12288 12288 0 100%
42% full, and not one more file can be created. You cannot add inodes to an existing ext4 filesystem: delete files, or re-create (re-format) it with more, using mkfs.ext4 -N count (mkfs = make filesystem). XFS, another Linux filesystem type, creates inodes as it needs them, which is one reason it is popular for exactly these workloads.
Finding where the inodes went is a du question with a different unit - --inodes counts files instead of bytes:
sudo du --inodes -x /srv/cache --max-depth=2 | sort -n | tail
2. A deleted file still held open. sudo lsof +L1
Someone ran rm on a big log that a process has open. The name is gone so du cannot see it, but the blocks are still allocated so df still counts them. du and df disagreeing by the size of one file is the tell.
$ sudo rm /data/app/debug.log
$ df -h /data
Filesystem Size Used Avail Use% Mounted on
/dev/vdb1 50G 42G 6.2G 87% /data
$ sudo du -sh /data/app
8.0K /data/app
$ sudo lsof +L1
COMMAND PID USER FD TYPE DEVICE SIZE/OFF NLINK NODE NAME
logwriter 1302 appuser 3w REG 253,17 42949672960 0 1157 /data/app/debug.log (deleted)
42G used according to df, 8K according to du, and lsof names the holder. Fixes, in order of preference:
- restart the service (it closes the old file and opens a fresh one),
- or empty the file through the descriptor, without restarting anything:
$ sudo truncate -s 0 /proc/1302/fd/3
$ df -h /data
Filesystem Size Used Avail Use% Mounted on
/dev/vdb1 50G 1.4G 47G 3% /data
truncate -s 0 FILE sets a file's size to 0 bytes. /proc/1302/fd/3 is the open file itself, so this empties the deleted inode. The process keeps its descriptor and keeps writing - into an empty file. Do not kill -9 it to free the space: a restart does the same job cleanly.
Then fix the actual cause: log rotation that the program knows about - logrotate's copytruncate option (copy the log, then truncate it in place, so the program's open file stays valid), or a signal that makes it reopen its log - or logging to the journal instead of a file.
3. Reserved blocks. sudo tune2fs -l /dev/vdb1 | grep -i reserved
ext4 keeps 5% of the filesystem as reserved blocks that only root may use. df subtracts them from Avail, so you see Avail 0 and Use% 100% while root can still write perfectly well. An ordinary user gets ENOSPC; root does not. That asymmetry is the diagnosis:
# from the 3am incident at the end of this chapter (/data/uploads is not there yet)
echo test > /data/uploads/test.txt
bash: echo: write error: No space left on device
sudo bash -c 'echo test > /data/uploads/root-test.txt' && echo "root can write"
root can write
$ sudo tune2fs -l /dev/vdb1 | grep -iE 'block count|reserved block'
Block count: 13107200
Reserved block count: 655360
(sudo bash -c '...' runs the whole command line, redirect included, as root - a plain sudo echo x > file would open the file as you; 1.11.) 655360 / 13107200 = 5%. The reserve exists so root can still log in and fix things, and so the filesystem has room to avoid fragmentation (files split into many scattered pieces). On a 50GB data volume, 5% is 2.5GB wasted for no benefit:
sudo tune2fs -m 1 /dev/vdb1 # -m = reserved percentage; 1% is plenty on a data volume
Leave it at 5% on the root filesystem.
Two more you will meet eventually
- A read-only remount. ext4 mounted with
errors=remount-roswitches itself to read-only after a disk error, to protect the data. The error isRead-only file system(EROFS), not ENOSPC - checkdmesg(the kernel's messages) andfindmnt -o TARGET,OPTIONS /. - Quotas. A quota is a per-user (or per-group) limit on disk use. Hitting it gives
Disk quota exceeded(EDQUOT) and has nothing to do with df.
The order matters
df -i first because it is instant and free. lsof +L1 second because it is the most common on a long-running box. tune2fs last because it only explains the case where root can write and a user cannot - which is a strong hint on its own. Say this order out loud in an interview, with the one-line reason for each, and you have answered the question.
What you can now do
- Name the three causes of ENOSPC on a disk with free space, in check order.
- Free a deleted-but-open file's space with a restart or
truncatevia /proc. - Read and change the root reserve with
tune2fs.