OnCallReady

Lesson 4.13 · Filesystem, Permissions, Disk · 16 min read

Links and inodes

In plain words

Imagine a library where each book has a shelf number, and the catalogue cards are separate from the book. You can have two catalogue cards for the same book, under two titles: both lead to the same copy. Throw one card away and the book is still reachable through the other. That is a hard link: two names, one inode.

A symbolic link is different: it is a sticky note saying "go look under this title". If that title's card is thrown away, the note leads nowhere; it dangles. And the librarian only removes the book from the shelf when no cards point to it and nobody is currently reading it. That last part explains why rm of a huge log that a process still has open frees no space: someone is still reading the book.

Why the filename is not the file

You delete a 40GB log and the disk does not get one byte back. A config symlink points at nothing after a deploy. mv to another disk takes ten minutes instead of an instant. All three make sense once you see that a name and a file are two separate things.

What you need to know already: 4.5 (ls -l, stat), 3.16 (file descriptors), 3.12 (lsof).

Inodes

A file is really an inode: a numbered record holding the metadata (owner, mode, timestamps, size) plus pointers to its data blocks - the chunks of disk (usually 4 KiB each) that hold the contents. A directory entry is just a name pointing at an inode number. The name is not stored in the inode at all.

$ echo data > a; ln a b; ln -s a c
$ ls -li a b c
3226 -rw-r--r-- 2 learner learner 5 Sep 22 20:00 a
3226 -rw-r--r-- 2 learner learner 5 Sep 22 20:00 b
3227 lrwxrwxrwx 1 learner learner 1 Sep 22 20:00 c -> a
 └┬─┘            └┬┘
  │               link count: TWO names point at inode 3226
  inode number: a and b are the same file; c is a different inode

ln TARGET NAME makes a hard link, ln -s TARGET NAME a symlink (both below). ls -i adds the inode number as the first column. The number after the mode is the link count: how many names point at this inode.

Hard link

ln a b creates a second name for the same inode - a hard link. Not a copy: change one and the other changes, because there is only one file. Delete a and b still works - the link count drops to 1 and the data stays.

There is no "original" and "link" after the fact. Both names are equally real; ls -li is the only way to tell they are the same file.

Limits, with the real errors:

$ ln b /data/x
ln: failed to create hard link '/data/x' => 'b': Invalid cross-device link
$ ln /etc /tmp/etclink
ln: /etc: hard link not allowed for directory

Inode numbers only mean something inside one filesystem (one formatted disk or partition; /data here is a separate disk - 4.18). So a hard link cannot cross into another one. Directories cannot be hard-linked because that would allow loops the kernel could never walk out of. (mv across filesystems works only because it silently becomes copy + delete - which is why it is slow.)

A directory's own link count is 2 + its number of subdirectories: its name in the parent, its own ., and each child's ...

Symbolic link

ln -s target name creates a symlink (symbolic link, soft link): a tiny file of its own that contains a path. It can cross filesystems, can point at directories, and can point at something that does not exist:

$ stat c
  File: c -> a
  Size: 1         	Blocks: 8          IO Block: 4096   symbolic link
$ stat -L c | head -3
  File: c
  Size: 5         	Blocks: 8          IO Block: 4096   regular file

The symlink's size is the length of the path it holds (1 byte: a). stat -L follows the link and describes the target instead.

Remove the target and the symlink dangles (points at nothing):

$ rm a
$ cat b
data
$ cat c
cat: c: No such file or directory
$ ls -l c
lrwxrwxrwx 1 learner learner 1 Sep 22 20:00 c -> a

b (the hard link) still has the data; c still exists but points at a name that is gone. ls is happy, cat is not. In colour, ls shows dangling links in red; find -xtype l lists them; readlink -f LINK prints the full path the link resolves to.

Relative targets are resolved from the link's directory, not from where you are standing: ln -s ../lib/app.jar /opt/app/bin/app.jar means /opt/app/lib/app.jar. This is why symlinks created with a relative path from the wrong directory dangle immediately.

The symlink's own permissions (lrwxrwxrwx) are meaningless - what counts is the target's, and x on the directories on the target's path (4.7). A symlink into a directory you cannot traverse is a symlink you cannot follow, however open the link looks.

The part that matters for disk space

The data blocks are freed when both of these reach zero:

  1. the number of names (the link count), and
  2. the number of open file descriptors on the inode.

So rm bigfile.log on a file that a running process still has open removes the name and frees nothing. We call it a deleted-but-open file: the inode lives on, invisible, until that process closes the fd or exits.

# after the rm of the 40G log in this chapter's disk mission (empty until then)
sudo lsof +L1
COMMAND    PID USER    FD TYPE DEVICE    SIZE/OFF NLINK NODE NAME
logwriter 1302 appuser 3w REG  253,17 42949672960     0 1157 /data/app/debug.log (deleted)

lsof +L1 means "open files with a link count less than 1". Columns: the program and its PID, the user, FD (3w = descriptor 3, open for writing), TYPE REG a regular file, SIZE/OFF its size in bytes, NLINK 0 no names left, NODE the inode number, and the old name marked (deleted). That is the whole diagnosis, and you will use it in a few steps.

Links in the wild

What you can now do

Why it helps

The inode model explains a whole family of real incidents: "I deleted the 40G log and the disk is still full" (a process holds it open, lsof +L1), "the config update did not apply" (the app followed a symlink once at startup), "our deploy symlink points nowhere" (a relative target resolved from the wrong directory), "mv to the other disk took forever" (it became copy plus delete).

It is also how common tools work: a current symlink pointing at the live release directory, rsnapshot backups with hard links, and /proc/PID/fd/N letting you recover or truncate a deleted file. Hard versus soft links is a very common junior interview question.

Commands in this lesson

echo ls touch ln stat rm cat

FAQ

Which one is the original after I create a hard link?

Neither. After ln a b, both names point to the same inode and are equally real; there is no link and original distinction in the filesystem. The link count in ls -l shows how many names exist, and ls -i shows the shared inode number. Deleting either name leaves the file intact until the count reaches zero and no process has it open.

Why can I not hard link across filesystems or to directories?

A hard link is a directory entry holding an inode number, and inode numbers are only unique within one filesystem, so a link to another filesystem's inode is meaningless. That gives "Invalid cross-device link". Directories cannot be hard-linked by users because it could create loops and multiple parents, which would break tools that walk the tree and the .. concept. Symlinks have neither limit.

How are relative symlink targets resolved?

Relative to the directory containing the link, not your current directory when you create or use it. ln -s ../lib/app.jar /opt/app/bin/app.jar points to /opt/app/lib/app.jar. Creating a relative link while thinking relative to your shell's location produces a dangling link. readlink -f link shows the fully resolved target, and find -xtype l lists broken links.

How can I get a deleted file back if a process still has it open?

Through /proc/PID/fd/N, which opens the still-existing inode. sudo cp /proc/1302/fd/3 /tmp/recovered.log copies the current contents. sudo lsof +L1 tells you which PID and fd. Once the process closes it or exits, the data is gone. The same path lets you free space without a restart using truncate -s 0 /proc/PID/fd/N.

Does mv keep hard links and inodes?

Within one filesystem, mv is a rename: the inode stays the same, with the same data, hard links, open file descriptors and ownership, and it is instant. Across filesystems, mv must copy the data to a new inode on the destination and delete the original, which takes time, changes the inode, breaks hard link relationships and does not affect processes already holding the old file open.

In an interview Junior

What is the difference between a hard link and a symbolic link?

A file is really an inode (metadata plus pointers to the data); a directory entry is just a name pointing at an inode number.

ls -li tells them apart: hard links share the inode number; a symlink is type l with -> target. And the data is only freed when the link count and the number of open file descriptors both reach zero - why rm of an open log frees nothing.

Also asked: You deleted a 40 GB log but disk usage did not change. Why? · Why does mv to another disk take minutes when mv on the same disk is instant? · How are relative symlink targets resolved?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.