When ps is missing, or lying
Every tool so far - ps, top, lsof - is a pretty printer over one folder: /proc. On a stripped-down server those tools may not be installed, and when two tools disagree you want the source. This lesson is a tour of it.
What you need to know already: 3.1 (PID, PPID, RSS, VSZ, threads), 3.12 (file descriptors, sockets), 2.21 (environment variables in units).
ps, top, lsof all read the same files
/proc is not on any disk. It is the kernel answering questions, formatted as files: when you cat one, the kernel writes the answer on the spot. Every running process has a directory named after its PID:
$ ls /proc/1210/
cgroup cmdline comm cwd environ exe fd limits oom_score
oom_score_adj stack stat status wchan
(The real kernel has about fifty entries per process; these are the ones worth knowing by name.) ps is a formatter over these files. When a tool is missing, cat still works.
cmdline: exactly how it was started
$ cat /proc/1210/cmdline
/usr/bin/java-Xmx512m-jar/opt/app/orders.jar
$ tr '\0' ' ' < /proc/1210/cmdline; echo
/usr/bin/java -Xmx512m -jar /opt/app/orders.jar
Arguments are separated by NUL bytes (the byte with value 0, which prints as nothing), not spaces, so a bare cat runs them together. tr '\0' ' ' (tr = translate characters: every NUL becomes a space) makes it readable; ; echo adds the missing newline.
The orders service is a Java program. -Xmx512m is the option that tells Java "keep at most 512 MB for the program's own data" (that area is called the heap; Chapter 5 comes back to it). cmdline is where you check which -Xmx a Java service really got, as opposed to what someone's config claims.
status: the human-readable summary
$ grep -E '^(Name|State|PPid|Uid|Threads|VmRSS|VmSize)' /proc/1210/status
Name: java
State: S (sleeping)
PPid: 1
Uid: 1001 1001 1001 1001
VmSize: 4200000 kB
VmRSS: 612000 kB
Threads: 64
(grep -E '^(A|B)' = lines that start with A or B.) Uid has four columns - real, effective, saved, filesystem - which only differ for special programs that switch users (Chapter 4). VmRSS is ps's RSS; VmSize is VSZ; VmHWM ("high water mark") is the peak RSS since start - useful after a memory spike has passed.
environ: the environment it runs with
$ cat /proc/1210/environ
cat: /proc/1210/environ: Permission denied
$ sudo cat /proc/1210/environ | tr '\0' '\n'
LANG=C.UTF-8
PATH=/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin:/snap/bin
INVOCATION_ID=7b3566fe71d4ac4d3a9b75698d79f7f8
JOURNAL_STREAM=8:21210
USER=appuser
HOME=/opt/app
Only the process's own user and root may read it (Chapter 4 explains how file permissions say that). Note the order: sudo tr ... < /proc/1210/environ would fail - the < is opened by your shell, as you, before sudo runs (the same trap as sudo echo > /etc/x, 1.11). INVOCATION_ID and JOURNAL_STREAM are how you can tell a process was started by systemd. And this is the second reason environment variables are a poor place for secrets (2.21 showed the first).
environ is the environment at the moment the program started. A program that changes its own environment later does not update this file.
exe and cwd: which binary, which directory
$ sudo ls -l /proc/1210/exe /proc/1210/cwd
lrwxrwxrwx 1 appuser appuser 1 Sep 22 20:00 /proc/1210/cwd -> /opt/app
lrwxrwxrwx 1 appuser appuser 43 Sep 22 20:00 /proc/1210/exe -> /usr/lib/jvm/java-21-openjdk-arm64/bin/java
Both are symlinks - entries that point at another path (the ->; Chapter 4 covers links). cwd is the current working directory; exe is the program file, followed through every symlink: /usr/bin/java is itself a link, and the kernel shows the real binary. If the binary was replaced by an upgrade after the process started, it reads ... (deleted): the running process is still the old version. That is how you find services that need a restart after patching.
fd: every open file, socket and pipe
$ sudo ls -l /proc/1210/fd | head -6
total 0
lrwx------ 1 appuser appuser 64 Sep 22 20:00 0 -> /dev/null
lrwx------ 1 appuser appuser 64 Sep 22 20:00 1 -> socket:[21210]
lrwx------ 1 appuser appuser 64 Sep 22 20:00 2 -> socket:[21210]
lrwx------ 1 appuser appuser 64 Sep 22 20:00 10 -> socket:[30010]
lr-x------ 1 appuser appuser 64 Sep 22 20:00 11 -> /opt/app/lib/orders-lib-1.jar
One symlink per file descriptor, named by its number:
- 0, 1, 2 are stdin, stdout, stderr. A systemd service's stdout is a socket to journald - that is how
echoin a service ends up in the journal. socket:[30010]- a network connection; the number is its ID, whichss -eandlsofcan match up.... (deleted)- a file removed from its folder but still held open. Chapter 4's 40G mystery lives here.
sudo ls /proc/PID/fd | wc -l is the fd count you compare with the Max open files line in /proc/PID/limits (3.16).
cgroup: who it belongs to
$ cat /proc/1210/cgroup
0::/system.slice/orders.service
$ cat /proc/$$/cgroup
0::/user.slice/user-1000.slice/session-1.scope
systemd puts every process into a cgroup (the kernel's named group of processes, which you gave limits to in 2.26). The file has one line, 0::<path>, and the path says who owns the process:
system.slice/X.service= a systemd service (a slice is systemd's folder of cgroups:system.slicefor services,user.slicefor logins).user.slice/.../session-N.scope= started from someone's login (a scope is a cgroup systemd made for processes it did not start itself).
One cat answers "who started this and who is supposed to manage it".
Later (Ch 10, 15): containers show up here too, as
docker-<id>.scope, and Kubernetes pods as/kubepods.slice/....
stat: the machine-readable line
$ cat /proc/1210/stat
1210 (java) S 1 1210 1210 0 -1 4194560 1200 0 0 0 12 4 ...
Field 1 PID, 2 (comm), 3 state, 4 PPID, 5 process group, 6 session (3.10), 14/15 user/system CPU time. Tools parse this; you read status.
wchan and stack: what it is waiting for
$ cat /proc/1210/wchan; echo
sk_wait_data
$ sudo cat /proc/1210/stack
wchan is the kernel function a sleeping task is parked in (3.3):
sk_wait_data= waiting for data on a socket.do_epoll_wait= an event loop waiting for any of its sockets (normal for nginx and most servers).rpc_wait_bit_killable= an NFS call (and killable).folio_wait_bit_common= waiting for writeback.io_schedule= a local disk.
stack (root only) is the whole kernel call chain.
The system-wide files
/proc/loadavg load averages, running/total, last PID
/proc/meminfo what free reads (Chapter 5)
/proc/cpuinfo one block per CPU (nproc counts them)
/proc/mounts what is mounted (Chapter 4)
/proc/sys/... kernel settings you can read and change
/proc/sys/vm/swappiness is the same setting as sysctl vm.swappiness (sysctl = the tool for kernel settings) - the dots become slashes.
What you can now do
- Answer "how was it started, as whom, with what, from where" from
/proc/<pid>/. - Know which files need root (
environ,fd/,exe,stack). - Tell a service from a hand-started process with one
cat /proc/<pid>/cgroup.