Up, not out of memory, and refusing everyone
A service can be running, have plenty of memory, and still reject every new connection - because it has hit its limit on open files. This lesson is how to spot that and raise the limit so it actually takes effect.
What you need to know already: 2.3 (drop-ins with systemctl edit), 3.12 (file descriptors, sockets).
Every socket is a file descriptor
An open file, a listening socket, every accepted connection, every pipe - each costs one fd. The kernel caps how many one process may hold. Run out and every new open fails at once with:
Too many open files (the error code is called EMFILE)
java.net.SocketException: Too many open files (how Java says it)
accept4() failed (24: Too many open files) (how nginx says it)
The service usually stays up, which is what makes it confusing: it is running, it has memory, and it cannot accept a single new connection.
Soft and hard limits
ulimit is a shell built-in that shows (and sets) the limits of the current shell and everything it starts; -n is the open-files limit:
$ ulimit -n 1024 the SOFT limit: what applies right now
$ ulimit -Hn 1048576 the HARD limit (-H): the ceiling you may raise it to
Any process can raise its soft limit up to the hard limit. Only root can raise the hard limit. The soft default stays at 1024 for a reason: an old system call, select(), cannot handle an fd numbered above 1023, so old programs would break. Programs that need more are expected to raise their own soft limit.
Java does exactly that: it raises its soft limit to the hard limit at startup, so a Java service under systemd normally shows 524288, not 1024. When a Java service is stuck at 1024, the hard limit is 1024 too - look for a LimitNOFILE=1024 someone set, which is what the orders unit on this box has.
Your shell's limits are not the service's limits. A service started by systemd gets its limits from systemd, not from you - so raising ulimit -n in your shell, or in /etc/security/limits.conf (the file that sets limits for logins), changes nothing for anything systemd starts. This is the most common reason "I already fixed that" is wrong.
Checking a running process
cat /proc/1210/limits | grep 'open files' the limit this process really has
sudo ls /proc/1210/fd | wc -l how many fds it has open RIGHT NOW
/proc/<pid>/limits is a table with the columns Limit, Soft Limit, Hard Limit, Units; the line you want is Max open files. /proc/<pid>/fd/ has one entry per open fd, so counting its lines with wc -l counts the fds.
The second command is the one that tells you whether you are about to hit the wall. 963 out of 1024 is not a warning sign, it is the incident starting.
Raising it properly
For a systemd service, in the unit or a drop-in:
[Service]
LimitNOFILE=65536
Then daemon-reload and restart the service - a limit is set when the process is created, so reloading systemd alone changes nothing about the process already running.
Verify where it counts, in /proc:
grep 'open files' /proc/$(systemctl show -p MainPID --value orders)/limits
What you can now do
- Read a process's real fd limit and current fd count from
/proc. - Explain soft vs hard, and why your shell's
ulimitdoes not reach a service. - Raise
LimitNOFILEwith a drop-in and a restart, and prove it took.