OnCallReady

Lesson 4.21 · Filesystem, Permissions, Disk · 15 min read

df shows free space but writes fail

In plain words

Imagine a car park that says "spaces free" on the sign, yet the barrier will not let you in. Three possible reasons. The car park has run out of parking tickets, even though bays are empty (inodes). A car was reported as gone, but it is still sitting in its bay because the owner never drove it out (a deleted file still held open). Or the last few bays are reserved for staff, so the sign counts them but visitors cannot use them (reserved blocks for root).

All three give the same message: No space left on device. Check them in order of speed: df -i for tickets, lsof +L1 for the ghost car, tune2fs -l for reserved bays, which is the case where root can park and you cannot.

Why "the disk has space" does not end the conversation

A write fails with No space left on device, and df -h says the disk is half empty. Unless you know the three ways that can happen - and which to check first - you lose an hour, or someone adds disk that changes nothing.

What you need to know already: 4.18 (df -h, df -i, tune2fs -l), 4.19 (du vs df), 4.13 (link count, deleted-but-open files), 3.14 (/proc/<pid>/fd).

Three causes, and the order to check them

All three produce the same error text, No space left on device. Its short name is ENOSPC - every kernel error has a code like this (called an errno), and the text is just its translation.

1. Inodes exhausted. df -i

Every file costs one inode, however small. A filesystem is created with a fixed number of them, and millions of tiny files - web session files, a cache directory, a mail queue - use up the inodes long before the blocks. This is inode exhaustion:

$ df -h /srv/cache        $ df -i /srv/cache
Size Used Avail Use%      Inodes IUsed IFree IUse%
2.0G 802M  1.2G  42%       12288 12288     0  100%

42% full, and not one more file can be created. You cannot add inodes to an existing ext4 filesystem: delete files, or re-create (re-format) it with more, using mkfs.ext4 -N count (mkfs = make filesystem). XFS, another Linux filesystem type, creates inodes as it needs them, which is one reason it is popular for exactly these workloads.

Finding where the inodes went is a du question with a different unit - --inodes counts files instead of bytes:

sudo du --inodes -x /srv/cache --max-depth=2 | sort -n | tail

2. A deleted file still held open. sudo lsof +L1

Someone ran rm on a big log that a process has open. The name is gone so du cannot see it, but the blocks are still allocated so df still counts them. du and df disagreeing by the size of one file is the tell.

$ sudo rm /data/app/debug.log
$ df -h /data
Filesystem  Size  Used Avail Use% Mounted on
/dev/vdb1    50G   42G  6.2G  87% /data
$ sudo du -sh /data/app
8.0K	/data/app
$ sudo lsof +L1
COMMAND    PID USER    FD TYPE DEVICE    SIZE/OFF NLINK NODE NAME
logwriter 1302 appuser 3w REG  253,17 42949672960     0 1157 /data/app/debug.log (deleted)

42G used according to df, 8K according to du, and lsof names the holder. Fixes, in order of preference:

$ sudo truncate -s 0 /proc/1302/fd/3
$ df -h /data
Filesystem  Size  Used Avail Use% Mounted on
/dev/vdb1    50G  1.4G   47G   3% /data

truncate -s 0 FILE sets a file's size to 0 bytes. /proc/1302/fd/3 is the open file itself, so this empties the deleted inode. The process keeps its descriptor and keeps writing - into an empty file. Do not kill -9 it to free the space: a restart does the same job cleanly.

Then fix the actual cause: log rotation that the program knows about - logrotate's copytruncate option (copy the log, then truncate it in place, so the program's open file stays valid), or a signal that makes it reopen its log - or logging to the journal instead of a file.

3. Reserved blocks. sudo tune2fs -l /dev/vdb1 | grep -i reserved

ext4 keeps 5% of the filesystem as reserved blocks that only root may use. df subtracts them from Avail, so you see Avail 0 and Use% 100% while root can still write perfectly well. An ordinary user gets ENOSPC; root does not. That asymmetry is the diagnosis:

# from the 3am incident at the end of this chapter (/data/uploads is not there yet)
echo test > /data/uploads/test.txt
bash: echo: write error: No space left on device
sudo bash -c 'echo test > /data/uploads/root-test.txt' && echo "root can write"
root can write
$ sudo tune2fs -l /dev/vdb1 | grep -iE 'block count|reserved block'
Block count:              13107200
Reserved block count:     655360

(sudo bash -c '...' runs the whole command line, redirect included, as root - a plain sudo echo x > file would open the file as you; 1.11.) 655360 / 13107200 = 5%. The reserve exists so root can still log in and fix things, and so the filesystem has room to avoid fragmentation (files split into many scattered pieces). On a 50GB data volume, 5% is 2.5GB wasted for no benefit:

sudo tune2fs -m 1 /dev/vdb1     # -m = reserved percentage; 1% is plenty on a data volume

Leave it at 5% on the root filesystem.

Two more you will meet eventually

The order matters

df -i first because it is instant and free. lsof +L1 second because it is the most common on a long-running box. tune2fs last because it only explains the case where root can write and a user cannot - which is a strong hint on its own. Say this order out loud in an interview, with the one-line reason for each, and you have answered the question.

What you can now do

Why it helps

This is one of the most common incidents and one of the most common interview questions for SRE roles, and the Notion checklist asks exactly this. Writes failing at 3am while dashboards show free space is confusing unless you have the three causes and their order in your head. Getting it right fast means an application recovers in minutes instead of someone adding disk that does not help.

Each cause maps to a real fix: purging small files and setting TTLs for inode exhaustion, restarting or truncating the holder for deleted-but-open logs, and tune2fs -m 1 on data volumes. The same reasoning separates the look-alikes - read-only remounts and quotas - which masquerade as disk problems but print different errors and need different fixes.

Commands in this lesson

df rm du lsof truncate tune2fs

FAQ

Can I add inodes to an ext4 filesystem?

Not directly. The inode count is set at mkfs time, by default roughly one per 16 KiB. Growing the filesystem with resize2fs adds inodes in proportion to the new space, but you cannot change the ratio. Options: delete files, grow the volume, or recreate it with mkfs.ext4 -N count or -i bytes-per-inode. For workloads with huge numbers of small files, XFS allocates inodes dynamically.

Why does root get more space than other users?

ext4 reserves 5% of blocks by default for UID 0 (configurable user and group). It exists so that system daemons and an admin can still work when users have filled the disk, and it helps the allocator avoid fragmentation. df hides the reserve from Avail. On the root filesystem, keep it; on large data volumes, 5% is a lot of wasted space, and tune2fs -m 1 is common.

Is truncating via /proc/PID/fd safe?

It is safe in the sense that the process keeps a valid descriptor and continues writing; you only lose the log contents written so far, which is usually the point. If the program does not open the file with O_APPEND, it keeps writing at its old offset, creating a sparse file whose apparent size looks large but uses few blocks. A service restart is cleaner when possible.

What is the difference between No space left, Read-only file system and Disk quota exceeded?

No space left on device (ENOSPC) means blocks or inodes ran out, or you hit the reserve. Read-only file system (EROFS) means the filesystem is mounted read-only, often because ext4 remounted it after I/O errors; check dmesg. Disk quota exceeded (EDQUOT) means a per-user, group or project quota was reached, independent of free space. Different errors, different fixes.

How do I stop this happening again?

Fix whatever produced the files, then watch both resources. Give caches and session directories a cleanup job (a systemd timer running find ... -mtime +N -delete), rotate logs with logrotate or leave them to the journal with a size limit, delete old backups on a schedule, and alert on inode usage (df -i) as well as space, with enough warning to act before 100%.

In an interview Junior

You deleted a 40 GB log file, but disk usage did not change. Why, and how do you get the space back?

rm removes the name, not the file. The data is freed only when the link count is 0 and no process has the file open - and a running service still has it open. So df still counts the blocks while du can no longer see the file: df and du disagreeing by the size of one file is the tell.

Find the holder: sudo lsof +L1 lists open files with no name left, e.g. logwriter 1302 appuser 3w ... /data/app/debug.log (deleted).

Get the space back, in order of preference: restart the service, so it closes the old file and opens a new one; or, without a restart, empty it through the descriptor: sudo truncate -s 0 /proc/1302/fd/3. Not kill -9. Then fix the cause: rotation the program knows about (logrotate's copytruncate, or a signal to reopen the log), or log to the journal.

Also asked: What is inode exhaustion, and how do you spot it? · Why does ext4 reserve 5% of blocks for root, and when would you lower it? · A user gets "No space left on device" but root can still write. What does that tell you?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.