OnCallReady

Lesson 3.16 · Processes & Signals · 12 min read

File descriptor limits

In plain words

Imagine a hotel receptionist who can hold only 1024 room keys at once. Every guest who checks in, every phone line and every open file takes a key. When all 1024 are out, the next guest is turned away at the door, even though the receptionist is awake and the hotel has empty rooms.

File descriptors are those keys, and ulimit -n is the number the receptionist can hold. The soft limit is the rule today; the hard limit is the most you are allowed to raise it to. The trap: a service started by systemd does not use your shell's rules. It uses its unit's LimitNOFILE=, and it only picks up a new value when it restarts. /proc/PID/limits shows the real number, and ls /proc/PID/fd | wc -l counts the keys in use.

Up, not out of memory, and refusing everyone

A service can be running, have plenty of memory, and still reject every new connection - because it has hit its limit on open files. This lesson is how to spot that and raise the limit so it actually takes effect.

What you need to know already: 2.3 (drop-ins with systemctl edit), 3.12 (file descriptors, sockets).

Every socket is a file descriptor

An open file, a listening socket, every accepted connection, every pipe - each costs one fd. The kernel caps how many one process may hold. Run out and every new open fails at once with:

Too many open files          (the error code is called EMFILE)
java.net.SocketException: Too many open files      (how Java says it)
accept4() failed (24: Too many open files)          (how nginx says it)

The service usually stays up, which is what makes it confusing: it is running, it has memory, and it cannot accept a single new connection.

Soft and hard limits

ulimit is a shell built-in that shows (and sets) the limits of the current shell and everything it starts; -n is the open-files limit:

$ ulimit -n        1024       the SOFT limit: what applies right now
$ ulimit -Hn       1048576    the HARD limit (-H): the ceiling you may raise it to

Any process can raise its soft limit up to the hard limit. Only root can raise the hard limit. The soft default stays at 1024 for a reason: an old system call, select(), cannot handle an fd numbered above 1023, so old programs would break. Programs that need more are expected to raise their own soft limit.

Java does exactly that: it raises its soft limit to the hard limit at startup, so a Java service under systemd normally shows 524288, not 1024. When a Java service is stuck at 1024, the hard limit is 1024 too - look for a LimitNOFILE=1024 someone set, which is what the orders unit on this box has.

Your shell's limits are not the service's limits. A service started by systemd gets its limits from systemd, not from you - so raising ulimit -n in your shell, or in /etc/security/limits.conf (the file that sets limits for logins), changes nothing for anything systemd starts. This is the most common reason "I already fixed that" is wrong.

Checking a running process

cat /proc/1210/limits | grep 'open files'     the limit this process really has
sudo ls /proc/1210/fd | wc -l                 how many fds it has open RIGHT NOW

/proc/<pid>/limits is a table with the columns Limit, Soft Limit, Hard Limit, Units; the line you want is Max open files. /proc/<pid>/fd/ has one entry per open fd, so counting its lines with wc -l counts the fds.

The second command is the one that tells you whether you are about to hit the wall. 963 out of 1024 is not a warning sign, it is the incident starting.

Raising it properly

For a systemd service, in the unit or a drop-in:

[Service]
LimitNOFILE=65536

Then daemon-reload and restart the service - a limit is set when the process is created, so reloading systemd alone changes nothing about the process already running.

Verify where it counts, in /proc:

grep 'open files' /proc/$(systemctl show -p MainPID --value orders)/limits

What you can now do

Why it helps

"Too many open files" is a classic production failure: a proxy or Java service stays up but refuses every new connection, often at peak traffic. The fix people try first, editing /etc/security/limits.conf or running ulimit -n in a shell, does nothing for systemd services, and knowing that saves an hour of "but I already raised it".

The right fix, a LimitNOFILE= drop-in plus a restart, verified in /proc/PID/limits, is a small, reviewable change. The same question appears in load tests, where the client machine runs out of descriptors before the server does. Monitoring fd usage against the limit catches leaks days before they turn into an outage.

Commands in this lesson

ulimit

FAQ

What is the difference between the soft and hard limit?

The soft limit is what is enforced right now. The hard limit is the ceiling: an unprivileged process can raise its soft limit up to the hard limit, and lower the hard limit, but only root (CAP_SYS_RESOURCE) can raise the hard limit. ulimit -n shows soft, ulimit -Hn hard. systemd's LimitNOFILE= can set both, as LimitNOFILE=65536 or soft:hard.

Why does limits.conf not affect my service?

/etc/security/limits.conf is applied by PAM (pam_limits) when a user logs in through a PAM service like ssh or su. Services started by systemd never go through PAM login, so they get systemd's defaults (DefaultLimitNOFILE in /etc/systemd/system.conf) or their unit's LimitNOFILE=. Set the limit in the unit or a drop-in, then restart.

Why is the default soft limit still 1024?

Compatibility. Old programs use the select() system call, which cannot handle descriptors numbered 1024 or higher, and a larger soft limit could let them silently break. So systemd keeps the soft limit at 1024 and a much larger hard limit (524288 in current systemd), and programs that can handle more simply raise their own soft limit. Many modern runtimes, including Java, do that automatically.

Is it safe to set LimitNOFILE very high?

High values like 65536 or even 1048576 are common for proxies and databases. The limit itself costs nothing, but it no longer protects you from a leak: a process leaking descriptors keeps consuming kernel memory until the system-wide fs.file-max or memory is affected. Also, some programs allocate tables sized to the limit at startup. Pick a value comfortably above peak use and monitor it.

How do I count a process's open files?

sudo ls /proc/PID/fd | wc -l is the quickest exact count. sudo lsof -p PID | wc -l gives a larger number because it also lists memory-mapped files, the cwd and other non-descriptor entries. Compare the fd count with the Max open files line in /proc/PID/limits. The system-wide total is in /proc/sys/fs/file-nr.

In an interview Junior

A service logs "Too many open files". How do you diagnose and fix it?

Every open file, socket and connection costs a file descriptor, and each process has a limit; at the limit every new open fails with EMFILE while the service stays "up".

Diagnose in /proc: grep 'open files' /proc/PID/limits for the real limit, sudo ls /proc/PID/fd | wc -l for how many are open right now. 963 out of 1024 is the incident starting.

Fix it in the unit, not in your shell: a drop-in with

[Service]
LimitNOFILE=65536

then daemon-reload and restart - limits are set when the process is created. Verify in /proc/$(systemctl show -p MainPID --value svc)/limits. The trap: ulimit -n in your shell or /etc/security/limits.conf changes nothing for a service systemd starts.

Also asked: What is a file descriptor? · What is the difference between the soft and the hard limit? · Why is the default soft limit still 1024?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.