Why this lesson
Half of all scripts read a file line by line: a list of hosts, a config, a CSV of users. The obvious loop silently eats backslashes, trims spaces, skips the last line, and - piped the wrong way - forgets every variable it set. A real incident at the end of this chapter (6.26) is exactly that last bug.
What you need to know already: while loops and conditions (6.3); IFS (6.1); < to feed a file to stdin (1.7); subshells and printf (6.5); quoting (6.6).
read, in one sentence
read VAR is a builtin that reads one line from stdin into VAR. It exits 0 if it got a line and non-zero at the end of the input - which is why it works as a while condition: the loop runs once per line and stops at the end.
The loop, and why every part of it is there
while IFS= read -r line; do
printf '%s\n' "$line"
done < input.txt
Four decisions in one line, each fixing a real bug:
read -r- without-r, backslashes are escape characters and vanish.IFS=(empty, and only for thisread) - without it, leading and trailing whitespace is stripped from every line.< input.txton thedone- the whole loop reads from the file, and runs in the current shell.printf '%s\n'rather thanecho- echo interprets-n,-eand (in some shells) backslashes in the data.
(IFS= read is the NAME=value command form: the variable is set for that one command only. Escape characters are characters with a special meaning: without -r, \t in the data is read as "a literal t" and the backslash is dropped.)
Watch each one fail:
$ printf 'C:\\temp\\new\n' > n; cat n
C:\temp\new
$ while read line; do echo "[$line]"; done < n
[C:tempnew]
$ while read -r line; do echo "[$line]"; done < n
[C:\temp\new]
$ printf ' indented \n' | while read -r l; do echo "[$l]"; done
[indented]
$ printf ' indented \n' | while IFS= read -r l; do echo "[$l]"; done
[ indented ]
shellcheck calls the missing -r SC2162.
Splitting a line into fields
A field is one piece of a line, and the delimiter is the character between the pieces. /etc/passwd has seven fields per line, delimited by : (name, password placeholder, UID, GID, comment, home, shell - 4.3).
Give read several names and it splits the line on IFS into fields; the last name gets everything left over:
$ while IFS=: read -r user _ uid gid _ home shell; do [ "$uid" -ge 1000 ] && echo "$user $uid $shell"; done < /etc/passwd
nobody 65534 /usr/sbin/nologin
learner 1000 /bin/bash
appuser 1001 /usr/sbin/nologin
_ is the conventional name for "a field I do not want". IFS=: applies to that one read only - the prefix-assignment form never changes IFS for the rest of the script.
$ IFS=: read -r a rest <<< 'x:y:z'; echo "$a | $rest"
x | y:z
$ IFS=, read -r a b c <<< '1,2 3,4'; echo "$a|$b|$c"
1|2 3|4
<<< is a here-string: the word, plus a newline, on stdin. Handy for splitting a variable without a pipe.
CSV (comma-separated values) is text with one record per line and fields separated by commas - what a spreadsheet exports. IFS=, read -r a b c handles simple CSV. A CSV with quoted fields that contain commas ("Popescu, Ana",platform) is beyond read; that needs a real CSV parser, not a shell loop.
The last line without a newline
read returns non-zero at end of file even if it read a partial line - so the final line of a file that does not end with a newline is silently skipped:
$ printf 'a\nb' > nl # no newline after b
$ while read -r l; do echo "[$l]"; done < nl
[a]
$ while read -r l || [[ -n $l ]]; do echo "[$l]"; done < nl
[a]
[b]
Files written by editors on Windows, by echo -n (echo without the final newline), or generated by other tools hit this constantly. The || [[ -n $l ]] idiom ("or, if something was read anyway") processes the fragment. ([[ -n $l ]] is true when $l is not empty - tests are in 6.10.)
Piping into while: the variables disappear
$ n=0
$ printf '1\n2\n3\n' | while read -r x; do n=$((n + x)); done
$ echo "$n"
0
Every part of a pipeline runs in its own subshell - a copy of the shell. The loop did add up to 6, in a copy that was thrown away when the pipeline ended. Feed the loop with a redirect instead, so it runs in the current shell:
$ n=0
$ while read -r x; do n=$((n + x)); done < <(printf '1\n2\n3\n')
$ echo "$n"
6
<(cmd) is process substitution: bash runs cmd and gives you a path (/dev/fd/63) to its output, and < <(cmd) redirects the loop's stdin from it. The same fix for a file: done < file, never cat file | while.
This is a real, common incident: a script reads its config with cat conf | while read k v; do ..., every setting is lost, and the defaults - often empty strings - are used instead.
find, filenames and NUL
Filenames can contain spaces and even newlines, so "one name per line" is not safe. The only safe separator is the NUL byte - the character with code 0, written \0 - which cannot appear in a path:
while IFS= read -r -d '' f; do
printf 'found: %s\n' "$f"
done < <(find /srv/releases -name '*.tmp' -print0)
find -print0 ends each path with NUL instead of a newline (4.15); read -d '' reads up to NUL instead of up to a newline (-d = delimiter).
mapfile: the whole file into an array
(Arrays - one variable holding a list - get their own lesson in 6.12.)
$ mapfile -t lines < /etc/hostname
$ echo "${#lines[@]}: ${lines[0]}"
1: oncall-lab
$ readarray -t hosts < <(printf '%s\n' web1 web2 'db 1')
$ printf '<%s>\n' "${hosts[@]}"
<web1>
<web2>
<db 1>
-t strips the trailing newline from each element (you almost always want it). readarray is the same builtin under another name. Use it when you need random access or a count; use while read when the input is large or streaming.
read from the terminal
read -r -p "Deploy to prod? [y/N] " answer
[[ $answer == [yY] ]] || exit 1
read -r -s -p "Password: " pw; echo
-p prints a prompt (to stderr), -s does not show what you type, -t 10 gives up after ten seconds. In a script that might run without a terminal (from cron or a systemd timer), check first: [[ -t 0 ]] is true only when stdin (fd 0) is a terminal.
while read, or something faster?
A shell loop runs its body once per line, in the shell - fine for tens of thousands of lines, slow for millions, and every external program in the body (date, grep) starts a new process per line. When the job is "compute over columns", a text tool like awk (Ch 7) is one process and far faster. Reach for while read when each line drives an action: a command, a file operation.
What you can now do
- Read a file line by line with
while IFS= read -r line; do ...; done < file. - Split lines into fields with
IFS=, read -r a b c, and keep the last line. - Keep a loop's variables by feeding it with
<or< <(cmd), never a pipe.