Processes & Signals: interview questions
The question you are most likely to get for each topic, a model answer, and what else comes up. From chapter 3 of the course.
What process states can you see in ps, and what does each mean? Junior
The first letter of the STAT column:
- R - running, or waiting in the queue for a CPU.
- S - interruptible sleep: waiting for something (a packet, a timer, input). Most processes, most of the time; healthy.
- D - uninterruptible sleep inside the kernel, usually waiting for a disk or a network file server. Signals are not delivered, so not even
kill -9works. - Z - zombie: already exited, waiting for its parent to collect the exit status.
- T - stopped: paused with Ctrl+Z or SIGSTOP.
- I - an idle kernel thread.
Modifiers follow: s session leader, l multi-threaded, + foreground, so Ssl is a normal service. Why it matters: load average counts R and D, so twelve D processes give a high load on an idle CPU. ps -eo pid,stat,comm lists them.
Also asked: Load is 15 on a 4-core box but the CPU is 90% idle. What is going on? · What is the difference between SIGTERM and SIGKILL? · What is a zombie process, and how do you get rid of it?
How do you find the processes using the most memory (or CPU) on a box? Junior
ps -eo pid,ppid,user,rss,vsz,stat,etime,comm --sort=-rss | head: -e every process, -o exactly these columns, --sort=-rss biggest first, head the top ten. Live, open top and press M to sort by memory or P by CPU.
Sort by RSS, not VSZ. RSS is the physical RAM the process really uses right now; VSZ is every address range it has reserved, including memory it never touched. A Java service using 600 MB of RAM routinely shows 4 GB of VSZ - meaningless.
The two old spellings: ps aux shows %CPU and %MEM ("what is eating the box?"), ps -ef shows the PPID ("who started this?"). A trailing = drops the header, for scripts: ps -o ppid= -p $$.
Also asked: What is the difference between ps aux and ps -ef? · What is the difference between RSS and VSZ? · How do you find the parent of a process?
Learn it: 3.1 ps, and the STAT column
Load average is 15 on a 4-core machine, but the CPU is 90% idle. What is going on? Junior
On Linux the load average counts processes that are R (running or waiting for a CPU) or D (stuck waiting inside the kernel, usually for I/O). With the CPU idle, it is not a CPU problem: about fifteen processes are blocked on something - an overloaded disk or a dead network file server (NFS). Adding CPUs changes nothing.
How I'd check: nproc and uptime to compare load with cores; top to confirm us/sy are near zero; then count the D processes - ps -eo pid,stat,comm or vmstat's b column. Then ask what they wait for: cat /proc/PID/wchan and sudo cat /proc/PID/stack; names like nfs_* and rpc_* mean NFS. Don't trust wa to decide - an NFS wait shows as plain idle. kill -9 won't clear them; the way out is the I/O completing, unmounting the dead share, or a reboot.
Also asked: A process will not die even with kill -9. What state is it in, and what does that tell you? · What do the three numbers of the load average mean? · What is the difference between load average and CPU utilisation?
Learn it: 3.3 Load average is not CPU usage
How do you tell quickly whether a box is CPU-bound, short on memory, or stuck on I/O? Junior
vmstat 1 5, ignoring the first line (it is the average since boot):
- CPU:
r(runnable processes) sustained above the number of cores andidlow. In top, highus/sy. - I/O:
bhigh - processes blocked in D state. That answers "is the load I/O?" directly. - Memory:
so(swap out) above zero, sustained. In top, read avail Mem, not "free".
Then top for the details: 1 for one line per core, P/M to sort, st for time a VM wanted a CPU but the hypervisor gave it to someone else. Remember %CPU is per core, so 180% on a 2-core box is almost all of it. Normalise load with nproc first, and for a ticket paste top -b -n 1 | head -15 rather than a screenshot.
Also asked: Explain the %Cpu(s) line in top. · What does st (steal) mean, and when would you see it? · Why does an NFS wait not show up in wa?
Learn it: 3.5 Reading top and vmstat, line by line
What is the difference between SIGTERM, SIGKILL, SIGHUP and SIGINT? Junior
- SIGTERM (15) - "please stop". The default of
kill, and what systemd sends first. Catchable: the program can finish requests, flush data and exit cleanly. - SIGKILL (9) - the kernel destroys the process. Cannot be caught, ignored or blocked; no cleanup at all.
- SIGHUP (1) - "hang up", originally "your terminal went away". By convention daemons reload their configuration on it:
sudo kill -HUP $(pgrep -o nginx)is whatsystemctl reload nginxdoes underneath. - SIGINT (2) - what Ctrl+C sends. Catchable.
The rule: TERM, wait, then KILL only if you must - kill -9 on a database means an unclean shutdown and a recovery on the next start. And neither reaches a process in D state. Find targets with pgrep -a before you pkill.
Also asked: How do you reload nginx's configuration without dropping connections? · Why should kill -9 not be your first resort? · How do you pause a process and resume it later?
Learn it: 3.6 Signals
What is a zombie process, and how do you get rid of it? Junior
A zombie (Z, <defunct> in ps) is a process that has already exited: the kernel freed its memory and files but keeps one small record, the exit status, until the parent collects it with wait() - that is called reaping. It uses no CPU and no memory, only a PID.
You cannot kill it - it is already dead, so kill -9 does nothing. The bug is in the parent, which is not reaping. Fix or stop the parent, and its zombies are handed to PID 1, which always reaps, so they vanish at once.
A few zombies are cosmetic; thousands mean a parent that will eventually use up every PID, and then nothing new can start. To find the parent: ps -eo pid,ppid,stat,comm, look for Z, and follow the PPID.
Also asked: What is an orphan process, and what happens to it? · What is special about PID 1? · Why does sudo kill -9 1 do nothing?
Learn it: 3.8 Zombies, orphans and PID 1
You need to run a two-hour migration over SSH. How do you make sure it survives a disconnect? Junior
Why it would die: when the connection drops, sshd closes the pseudo-terminal, the kernel sends SIGHUP to the session leader (your bash), bash re-sends SIGHUP to every job, and SIGHUP's default action is to terminate.
The fixes, weakest to strongest:
nohup ./migrate.sh > migrate.log 2>&1 &- starts it with SIGHUP ignored. Nothing else: no logs, no supervision.setsid- a new session with no terminal, so no hangup at all.- tmux - the shell lives in the tmux server, not your SSH session; detach with
Ctrl+b d, reconnect andtmux attachlater, output still on screen. sudo systemd-run --unit=migrate /usr/local/bin/migrate.sh- a real transient unit: its own cgroup, output injournalctl -u migrate,systemctl stopwhen you want it gone.
Also asked: What happens, signal by signal, when an SSH connection drops? · What is the difference between a process group and a session? · How do you find processes left behind by dead SSH sessions?
A deploy fails with "Address already in use" on port 8080. How do you find what holds the port? Junior
sudo ss -tlnp (TCP, listening, numeric, with the process) or sudo lsof -i :8080. The sudo matters: without root both only see your own processes, so the port looks free and you conclude nothing is there.
Then turn the PID into an owner: systemctl status PID names the unit whose cgroup holds it - often the old instance that did not stop. If it has no unit, it was started by hand and nobody supervises it; ps -o pid,ppid,etime,args -p PID shows how long it has been there.
Stop the right owner properly (systemctl stop), not kill -9 on whatever lsof -t returns. And if a process is up but does nothing, sudo strace -f -p PID shows what it is blocked on: a read() that never returns is waiting on a socket with no timeout.
Also asked: A service is running but requests hang. How would you see what it is waiting for? · How do you see which files a process has open? · Why is strace risky on a busy production service?
Learn it: 3.12 lsof and strace
What useful information can you find under /proc/PID? Junior
/proc is the kernel answering questions as files; ps, top and lsof are formatters over it, so it works when they are missing.
cmdline- exactly how it was started (NUL-separated:tr '\0' ' '), e.g. which-Xmxa Java service really got.status- name, state, PPID, UIDs, threads,VmRSS(= RSS),VmHWM(peak RSS).environ- its environment at start (root or the owner only).exeandcwd- the real binary and working directory;exe ... (deleted)means it still runs an old version after an upgrade.fd/- one symlink per open file, socket or pipe.limits- e.g.Max open files.cgroup- who owns it:system.slice/X.serviceor a login'ssession-N.scope.wchanandstack- what it is waiting for in the kernel.
Also asked: A box has no ps or lsof and you may not install anything. How do you inspect its processes? · How can you tell whether a process was started by systemd or by hand? · After a library upgrade, how do you find processes still running the old version?
Learn it: 3.14 /proc: every process is a directory
A service logs "Too many open files". How do you diagnose and fix it? Junior
Every open file, socket and connection costs a file descriptor, and each process has a limit; at the limit every new open fails with EMFILE while the service stays "up".
Diagnose in /proc: grep 'open files' /proc/PID/limits for the real limit, sudo ls /proc/PID/fd | wc -l for how many are open right now. 963 out of 1024 is the incident starting.
Fix it in the unit, not in your shell: a drop-in with
[Service]
LimitNOFILE=65536
then daemon-reload and restart - limits are set when the process is created. Verify in /proc/$(systemctl show -p MainPID --value svc)/limits. The trap: ulimit -n in your shell or /etc/security/limits.conf changes nothing for a service systemd starts.
Also asked: What is a file descriptor? · What is the difference between the soft and the hard limit? · Why is the default soft limit still 1024?
Learn it: 3.16 File descriptor limits
Users see a few errors on every deploy. What is probably going on? Junior
Something is wrong with the graceful stop. Every supervisor stops a service the same way: an optional stop hook (ExecStop=), SIGTERM, a grace period (TimeoutStopSec=, 90 s by default), then SIGKILL. A well-behaved program uses the grace period to drain: stop taking new work, finish the requests in flight, exit.
Two usual failures:
- It ignores SIGTERM - no handler, or a wrapper script that never passes the signal on. Every stop takes the full grace period and ends in SIGKILL, cutting requests off. The tell: it always exits 137 instead of 143. Fix: handle SIGTERM, and end wrapper scripts with
exec ./the-real-program. - Traffic still arrives after SIGTERM - the load balancer has not stopped routing to it yet, and it closes its socket at once: "connection refused". Fix: keep serving a few seconds after SIGTERM (or
sleep 5in the stop hook).
Also asked: What do exit codes 143, 137 and 130 tell you about how a process ended? · What is a grace period, and how would you choose its length? · Why does a shell script wrapper break a graceful stop, and what does exec fix?
Learn it: 3.18 A graceful stop: SIGTERM, grace period, SIGKILL
Practise these answers with flashcards and labs Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.