OnCallReady

Lesson 3.10 · Processes & Signals · 14 min read

Sessions, SIGHUP, and surviving a disconnect

In plain words

Imagine phoning a friend and asking them to read a two-hour story aloud to the children in the room. If the phone line drops, the old rule is: when the call ends, everyone who was told "keep going while I am on the line" stops. The line dropping is the signal SIGHUP, "hang up".

When your SSH connection dies, the terminal closes, the kernel sends SIGHUP to your shell, and bash passes it to every job it started, so your migration dies. nohup is telling the reader "ignore the hang-up". setsid gives them their own phone that no one can hang up. tmux keeps the call going in another room that you can walk back into. systemd-run hands the story to the building manager, who logs it and keeps it going properly.

The question

You started a two-hour data migration over SSH. Your Wi-Fi dropped after twenty minutes. Is the migration still running?

Almost certainly no - and the reason is one signal. This lesson is why, and the four ways to start long work so it survives.

What you need to know already: 3.2 (your shell hangs off sshd-session), 3.6 (signals, jobs, &, %1), 3.8 (orphans go to PID 1), 2.26 (systemd-run).

Sessions and process groups

Two groupings the kernel keeps for every process:

A session can have one controlling terminal. Over SSH that is a pseudo-terminal (pty): a fake terminal device that sshd creates for your connection, shown as pts/0, pts/1...

$ sleep 600 &
[1] 17401
$ ps -o pid,ppid,pgid,sid,tty,stat,cmd --ppid $$
    PID    PPID    PGID     SID TT       STAT CMD
  17401    1566   17401    1566 pts/0    S    sleep 600

(--ppid $$ = only processes whose parent is your shell.)

ps -o sid= on your shell prints its own PID: a login shell leads its session.

What a disconnect actually does

  1. The network drops. sshd notices (a keepalive - a periodic "are you still there?" message - goes unanswered, or a write fails) and closes the pseudo-terminal pts/0.
  2. The kernel sends SIGHUP ("hang up" - from modem days) to the session leader: your bash.
  3. An interactive bash, on receiving SIGHUP, re-sends SIGHUP to every job in its job table, running or stopped, and then exits.
  4. SIGHUP's default action is to terminate. Your migration dies.

A foreground command dies the same way: it is in the terminal's foreground group, which gets the hangup directly.

kill -9 $$ from 3.8 is different: bash is killed before it can forward anything, so its background jobs are left behind as orphans (re-parented to PID 1). They keep running until they touch the terminal that no longer exists: a read or write then fails with an I/O error (EIO), which most programs treat as fatal - and a stopped job in an orphaned group gets SIGHUP + SIGCONT from the kernel and dies at once.

The fixes, from weakest to strongest

nohup - start the command with SIGHUP ignored:

$ nohup ./migrate.sh &
[1] 17420
nohup: ignoring input and appending output to 'nohup.out'

"Ignored" is remembered by the new program when it starts, so the hangup is simply dropped. Output that would go to the terminal goes to ./nohup.out (or ~/nohup.out if the current directory is not writable). Redirect it yourself to choose: nohup ./migrate.sh > migrate.log 2>&1 &.

nohup does not protect against SIGTERM, SIGKILL, running out of memory, or a reboot.

disown - remove a job from bash's job table after starting it:

$ ./migrate.sh &
$ disown %1          # bash forgets it: no SIGHUP will be forwarded
$ disown -h %1       # keep it in the table, but mark it: skip it on hangup

Useful when you already started something the naive way. It still has the dead terminal as stdout, so anything it prints after the disconnect can kill it (SIGPIPE - "the other end of your output is gone" - or EIO). Redirect output when you start it.

setsid - run it in a brand-new session with no controlling terminal:

$ setsid ./migrate.sh > migrate.log 2>&1 < /dev/null &

There is no terminal to hang up, so there is no SIGHUP at all.

tmux / screen - the one you actually use interactively. tmux ("terminal multiplexer") runs your shell inside a tmux server process that is not part of your SSH session. You detach (Ctrl+b d), disconnect, reconnect later, run tmux attach, and the migration is still on screen with its output. Install it on any box you look after.

systemd-run - the one you use for anything that matters:

$ sudo systemd-run --unit=export-test sleep 300
Running as unit: export-test.service; invocation ID: 12febe065567547b12bc0a9d43a44c6d
$ systemctl status export-test
$ journalctl -u export-test -f

(--unit=NAME names the unit. On a real box the command is your script, systemd-run --unit=nightly-export /usr/local/bin/export.sh; sleep stands in for it here. systemd-run checks the path before it starts anything: a typo fails at once with Failed to find executable ..., not later as a 203/EXEC.)

It becomes a real transient unit - a service that exists only until it finishes, with no unit file: its own cgroup under system.slice, no terminal, output in the journal, resource limits with -p MemoryMax=..., and systemctl stop when you want it gone. Your SSH session can die ten times; the unit does not care.

The flip side: the orphan nobody owns

Every one of these techniques creates a process that outlives its creator. The chapter's incident - a CPU burner with PPID 1 and no unit - is what that looks like a week later, when nobody remembers starting it. systemd-run at least leaves a named unit and a journal behind. nohup leaves nothing but a PID.

Checking what survived

$ ps -eo pid,ppid,sid,tty,stat,etime,cmd | awk '$2 == 1 && $4 != "?"'

(awk '$2 == 1 && $4 != "?"' keeps lines whose 2nd column, PPID, is 1 and whose 4th, TTY, is not ?.) Processes re-parented to PID 1 that still claim a terminal are leftovers from a dead session. systemctl status <PID> names the session scope they came from.

What you can now do

Why it helps

Losing a long-running job to a Wi-Fi drop happens to everyone once: a database migration, a large rsync, a backfill, a big upgrade script. Knowing SIGHUP is the cause, and reaching for tmux or systemd-run before starting, prevents half-finished migrations that are painful to recover from.

The flip side is operational hygiene: every nohup job is an unsupervised process that someone will later find with PPID 1, no unit and no logs, like this chapter's CPU-burner incident. Using systemd-run --unit=name leaves a named unit, journal logs and resource limits, which is what you want on any shared server. This is also a common "how would you run a long task on a remote server" interview question.

Commands in this lesson

ps sleep nohup disown setsid systemd-run systemctl journalctl

FAQ

Is nohup enough to keep a job running after I log out?

Usually, for the hang-up itself: nohup starts the command with SIGHUP ignored, so the disconnect does not kill it. It does not protect against SIGTERM, the OOM killer or a reboot, and some logout configurations still end it: systemd-logind with KillUserProcesses=yes kills everything in the session scope. Ubuntu leaves that setting off by default. Always redirect output yourself, or it goes to nohup.out.

What is the difference between disown and nohup?

nohup sets SIGHUP to ignored before the program starts, and redirects output away from the terminal. disown works afterwards: it removes a job from bash's job table (or with -h marks it) so bash does not forward SIGHUP when it exits. disown does not change the program's own terminal output, so if it writes to the closed terminal later, it can die with SIGPIPE or an I/O error.

Why use tmux instead of nohup?

tmux keeps a whole interactive shell alive in a server process independent of your SSH connection. You can detach, reconnect from anywhere, see the live output, answer prompts and scroll back. nohup gives you a blind background process and a log file. For interactive or long operational work (upgrades, migrations, debugging sessions), tmux or screen is the standard tool; install it on every box you administer.

Why does Ctrl+C not kill my background jobs?

Ctrl+C makes the terminal driver send SIGINT to the foreground process group only. bash puts each job in its own process group, and background jobs are not the foreground group, so they are not affected. The same is true for Ctrl+Z. To signal a background job, use kill %1 or bring it forward with fg first.

What does systemd-run give me over running the script in tmux?

A real transient unit: a name, its own cgroup, output in the journal (journalctl -u name), status with systemctl status name, clean stopping with systemctl stop, and resource limits with -p MemoryMax= or -p CPUQuota=. It survives your session entirely and other admins can see what it is. tmux is better for interactive work; systemd-run is better for unattended jobs.

In an interview Junior

You need to run a two-hour migration over SSH. How do you make sure it survives a disconnect?

Why it would die: when the connection drops, sshd closes the pseudo-terminal, the kernel sends SIGHUP to the session leader (your bash), bash re-sends SIGHUP to every job, and SIGHUP's default action is to terminate.

The fixes, weakest to strongest:

Also asked: What happens, signal by signal, when an SSH connection drops? · What is the difference between a process group and a session? · How do you find processes left behind by dead SSH sessions?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.