The question
You started a two-hour data migration over SSH. Your Wi-Fi dropped after twenty minutes. Is the migration still running?
Almost certainly no - and the reason is one signal. This lesson is why, and the four ways to start long work so it survives.
What you need to know already: 3.2 (your shell hangs off sshd-session), 3.6 (signals, jobs, &, %1), 3.8 (orphans go to PID 1), 2.26 (systemd-run).
Sessions and process groups
Two groupings the kernel keeps for every process:
- A process group is one job: a command, or a whole pipeline. Its ID is the PGID.
- A session is everything started from one login. Its ID, the SID, is the PID of the session leader - your login shell.
A session can have one controlling terminal. Over SSH that is a pseudo-terminal (pty): a fake terminal device that sshd creates for your connection, shown as pts/0, pts/1...
$ sleep 600 &
[1] 17401
$ ps -o pid,ppid,pgid,sid,tty,stat,cmd --ppid $$
PID PPID PGID SID TT STAT CMD
17401 1566 17401 1566 pts/0 S sleep 600
(--ppid $$ = only processes whose parent is your shell.)
- SID 1566 - the session is led by your login shell. Everything you start from it is in that session, attached to its terminal
pts/0. - PGID 17401 - bash put the job in its own process group. Ctrl+C and Ctrl+Z are delivered to the foreground group only, which is why Ctrl+C does not kill your background jobs.
ps -o sid= on your shell prints its own PID: a login shell leads its session.
What a disconnect actually does
- The network drops. sshd notices (a keepalive - a periodic "are you still there?" message - goes unanswered, or a write fails) and closes the pseudo-terminal
pts/0. - The kernel sends SIGHUP ("hang up" - from modem days) to the session leader: your bash.
- An interactive bash, on receiving SIGHUP, re-sends SIGHUP to every job in its job table, running or stopped, and then exits.
- SIGHUP's default action is to terminate. Your migration dies.
A foreground command dies the same way: it is in the terminal's foreground group, which gets the hangup directly.
kill -9 $$ from 3.8 is different: bash is killed before it can forward anything, so its background jobs are left behind as orphans (re-parented to PID 1). They keep running until they touch the terminal that no longer exists: a read or write then fails with an I/O error (EIO), which most programs treat as fatal - and a stopped job in an orphaned group gets SIGHUP + SIGCONT from the kernel and dies at once.
The fixes, from weakest to strongest
nohup - start the command with SIGHUP ignored:
$ nohup ./migrate.sh &
[1] 17420
nohup: ignoring input and appending output to 'nohup.out'
"Ignored" is remembered by the new program when it starts, so the hangup is simply dropped. Output that would go to the terminal goes to ./nohup.out (or ~/nohup.out if the current directory is not writable). Redirect it yourself to choose: nohup ./migrate.sh > migrate.log 2>&1 &.
nohup does not protect against SIGTERM, SIGKILL, running out of memory, or a reboot.
disown - remove a job from bash's job table after starting it:
$ ./migrate.sh &
$ disown %1 # bash forgets it: no SIGHUP will be forwarded
$ disown -h %1 # keep it in the table, but mark it: skip it on hangup
Useful when you already started something the naive way. It still has the dead terminal as stdout, so anything it prints after the disconnect can kill it (SIGPIPE - "the other end of your output is gone" - or EIO). Redirect output when you start it.
setsid - run it in a brand-new session with no controlling terminal:
$ setsid ./migrate.sh > migrate.log 2>&1 < /dev/null &
There is no terminal to hang up, so there is no SIGHUP at all.
tmux / screen - the one you actually use interactively. tmux ("terminal multiplexer") runs your shell inside a tmux server process that is not part of your SSH session. You detach (Ctrl+b d), disconnect, reconnect later, run tmux attach, and the migration is still on screen with its output. Install it on any box you look after.
systemd-run - the one you use for anything that matters:
$ sudo systemd-run --unit=export-test sleep 300
Running as unit: export-test.service; invocation ID: 12febe065567547b12bc0a9d43a44c6d
$ systemctl status export-test
$ journalctl -u export-test -f
(--unit=NAME names the unit. On a real box the command is your script, systemd-run --unit=nightly-export /usr/local/bin/export.sh; sleep stands in for it here. systemd-run checks the path before it starts anything: a typo fails at once with Failed to find executable ..., not later as a 203/EXEC.)
It becomes a real transient unit - a service that exists only until it finishes, with no unit file: its own cgroup under system.slice, no terminal, output in the journal, resource limits with -p MemoryMax=..., and systemctl stop when you want it gone. Your SSH session can die ten times; the unit does not care.
The flip side: the orphan nobody owns
Every one of these techniques creates a process that outlives its creator. The chapter's incident - a CPU burner with PPID 1 and no unit - is what that looks like a week later, when nobody remembers starting it. systemd-run at least leaves a named unit and a journal behind. nohup leaves nothing but a PID.
Checking what survived
$ ps -eo pid,ppid,sid,tty,stat,etime,cmd | awk '$2 == 1 && $4 != "?"'
(awk '$2 == 1 && $4 != "?"' keeps lines whose 2nd column, PPID, is 1 and whose 4th, TTY, is not ?.) Processes re-parented to PID 1 that still claim a terminal are leftovers from a dead session. systemctl status <PID> names the session scope they came from.
What you can now do
- Explain exactly why a dropped SSH connection kills your foreground work.
- Start long work so it survives:
nohup,setsid,tmux,systemd-run. - Find leftovers from dead sessions (PPID 1, a terminal, a session scope).