OnCallReady

Chapter 3 Processes & Signals

ps and what STAT means, load average that is not CPU, signals, zombies, lsof and strace, fd limits, and what a graceful stop really is.

In plain words

Imagine a busy kitchen. Every cook has a ticket number, a boss who hired them, a station, and a state: cooking, waiting for the oven, taking a break, or already gone home but still on the rota because nobody crossed them off. The head chef can shout instructions at any cook: "finish up and leave", "stop right now", "re-read the menu".

Processes are the cooks. ps and top read the rota: PID, parent, state (R S D Z T). Signals are the shouts: SIGTERM, SIGKILL, SIGHUP. /proc is each cook's personnel file, lsof lists what each one is holding, strace watches their hands. This chapter is how you tell a slow kitchen from a stuck one, and how to send someone home without dropping the plates.

Why it matters on call

Most production symptoms show up as process behaviour first: a box that is "slow", a load average of 15, a service that will not die, a stop that always takes 90 seconds, "Too many open files", a CPU burner nobody owns. This chapter gives you the reading skills to go from symptom to cause in minutes: which state, which parent, which unit, what it is waiting for.

It is also the chapter behind a lot of interview classics: "a process will not die with kill -9", "load is high but CPU is idle", "what happens between a stop request and the process disappearing". And it looks at systemd from the other side: PID 1 duties, the SIGTERM-grace-SIGKILL sequence and fd limits are mechanisms every later tool you meet is built on.

Lessons

  1. ps, and the STAT column
  2. Load average is not CPU usage
  3. Reading top and vmstat, line by line
  4. Signals
  5. Zombies, orphans and PID 1
  6. Sessions, SIGHUP, and surviving a disconnect
  7. lsof and strace
  8. /proc: every process is a directory
  9. File descriptor limits
  10. A graceful stop: SIGTERM, grace period, SIGKILL

13 hands-on labs (missions, incidents and drills) run in the terminal: Open this chapter in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.

Questions people ask

What is the difference between a process and a thread?

A process is one running program with its own memory, its own table of open files and its own PID. Threads are separate lines of work inside one process that share its memory and open files. On Linux the kernel schedules both the same way; each thread has its own ID, visible in /proc/PID/task/ and with ps -L or top -H. A Java service shows as one process with dozens of threads. Load average and /proc/loadavg count threads.

Is kill only for killing?

No. kill sends any signal, and most signals are messages: SIGHUP asks daemons to reload, SIGSTOP and SIGCONT pause and resume, SIGUSR1 and SIGUSR2 mean whatever the program defines (reopen the log file, print a status report). Only SIGKILL is guaranteed to end a process. The name is historical; think of it as "send signal".

Is top on Linux the same as Activity Monitor or top on my Mac?

Same idea, different tool. macOS top is a BSD variant with other columns and keys. Linux top comes from the procps-ng package and reads /proc: it shows states like D, the wa and st CPU fields, and avail Mem. Load average on Linux also counts D-state processes, unlike most other Unix systems. Learn the Linux version, because that is what servers run.

Do I need this if systemd restarts things for me?

Yes. Restart= brings a crashed service back, but it cannot tell you why a process is stuck in D state, leaking file descriptors, ignoring SIGTERM or burning CPU. Those questions are answered with ps, /proc, lsof and strace. Restarts also hide root causes: a service in an endless activating (auto-restart) loop looks alive on a dashboard. Knowing the mechanism is how you find the bug behind the loop.

Why does this chapter come after systemd?

Because systemd is how most processes on a server get started, supervised and stopped, and it already introduced signals, cgroups and exit codes from the service side. This chapter looks at the same things from the process side: states, parents, signals and limits. With both views you can map any PID back to its unit (systemctl status PID) and understand what a stop really does.