OnCallReady

Lesson 6.8 · Bash Scripting · 22 min read

Reading input: while read, IFS and mapfile

In plain words

Imagine reading a list aloud, one line at a time. A careless reader skips the backslashes, trims the spaces at the start and end of each line, and stops before the last line if it has no full stop. A careful reader says every character exactly as written.

while IFS= read -r line; do ...; done < file is the careful reader: -r keeps backslashes, IFS= keeps leading and trailing spaces, and < file on the done keeps the loop in the same room so its notes survive. Piping into while sends the reader into another room, and everything they counted is lost when they come back. mapfile -t reads the whole list into an array at once. IFS=: read -r user _ uid splits one line into named fields.

Why this lesson

Half of all scripts read a file line by line: a list of hosts, a config, a CSV of users. The obvious loop silently eats backslashes, trims spaces, skips the last line, and - piped the wrong way - forgets every variable it set. A real incident at the end of this chapter (6.26) is exactly that last bug.

What you need to know already: while loops and conditions (6.3); IFS (6.1); < to feed a file to stdin (1.7); subshells and printf (6.5); quoting (6.6).

read, in one sentence

read VAR is a builtin that reads one line from stdin into VAR. It exits 0 if it got a line and non-zero at the end of the input - which is why it works as a while condition: the loop runs once per line and stops at the end.

The loop, and why every part of it is there

while IFS= read -r line; do
  printf '%s\n' "$line"
done < input.txt

Four decisions in one line, each fixing a real bug:

(IFS= read is the NAME=value command form: the variable is set for that one command only. Escape characters are characters with a special meaning: without -r, \t in the data is read as "a literal t" and the backslash is dropped.)

Watch each one fail:

$ printf 'C:\\temp\\new\n' > n; cat n
C:\temp\new
$ while read line; do echo "[$line]"; done < n
[C:tempnew]
$ while read -r line; do echo "[$line]"; done < n
[C:\temp\new]

$ printf '  indented  \n' | while read -r l; do echo "[$l]"; done
[indented]
$ printf '  indented  \n' | while IFS= read -r l; do echo "[$l]"; done
[  indented  ]

shellcheck calls the missing -r SC2162.

Splitting a line into fields

A field is one piece of a line, and the delimiter is the character between the pieces. /etc/passwd has seven fields per line, delimited by : (name, password placeholder, UID, GID, comment, home, shell - 4.3).

Give read several names and it splits the line on IFS into fields; the last name gets everything left over:

$ while IFS=: read -r user _ uid gid _ home shell; do [ "$uid" -ge 1000 ] && echo "$user $uid $shell"; done < /etc/passwd
nobody 65534 /usr/sbin/nologin
learner 1000 /bin/bash
appuser 1001 /usr/sbin/nologin

_ is the conventional name for "a field I do not want". IFS=: applies to that one read only - the prefix-assignment form never changes IFS for the rest of the script.

$ IFS=: read -r a rest <<< 'x:y:z'; echo "$a | $rest"
x | y:z
$ IFS=, read -r a b c <<< '1,2 3,4'; echo "$a|$b|$c"
1|2 3|4

<<< is a here-string: the word, plus a newline, on stdin. Handy for splitting a variable without a pipe.

CSV (comma-separated values) is text with one record per line and fields separated by commas - what a spreadsheet exports. IFS=, read -r a b c handles simple CSV. A CSV with quoted fields that contain commas ("Popescu, Ana",platform) is beyond read; that needs a real CSV parser, not a shell loop.

The last line without a newline

read returns non-zero at end of file even if it read a partial line - so the final line of a file that does not end with a newline is silently skipped:

$ printf 'a\nb' > nl                      # no newline after b
$ while read -r l; do echo "[$l]"; done < nl
[a]
$ while read -r l || [[ -n $l ]]; do echo "[$l]"; done < nl
[a]
[b]

Files written by editors on Windows, by echo -n (echo without the final newline), or generated by other tools hit this constantly. The || [[ -n $l ]] idiom ("or, if something was read anyway") processes the fragment. ([[ -n $l ]] is true when $l is not empty - tests are in 6.10.)

Piping into while: the variables disappear

$ n=0
$ printf '1\n2\n3\n' | while read -r x; do n=$((n + x)); done
$ echo "$n"
0

Every part of a pipeline runs in its own subshell - a copy of the shell. The loop did add up to 6, in a copy that was thrown away when the pipeline ended. Feed the loop with a redirect instead, so it runs in the current shell:

$ n=0
$ while read -r x; do n=$((n + x)); done < <(printf '1\n2\n3\n')
$ echo "$n"
6

<(cmd) is process substitution: bash runs cmd and gives you a path (/dev/fd/63) to its output, and < <(cmd) redirects the loop's stdin from it. The same fix for a file: done < file, never cat file | while.

This is a real, common incident: a script reads its config with cat conf | while read k v; do ..., every setting is lost, and the defaults - often empty strings - are used instead.

find, filenames and NUL

Filenames can contain spaces and even newlines, so "one name per line" is not safe. The only safe separator is the NUL byte - the character with code 0, written \0 - which cannot appear in a path:

while IFS= read -r -d '' f; do
  printf 'found: %s\n' "$f"
done < <(find /srv/releases -name '*.tmp' -print0)

find -print0 ends each path with NUL instead of a newline (4.15); read -d '' reads up to NUL instead of up to a newline (-d = delimiter).

mapfile: the whole file into an array

(Arrays - one variable holding a list - get their own lesson in 6.12.)

$ mapfile -t lines < /etc/hostname
$ echo "${#lines[@]}: ${lines[0]}"
1: oncall-lab
$ readarray -t hosts < <(printf '%s\n' web1 web2 'db 1')
$ printf '<%s>\n' "${hosts[@]}"
<web1>
<web2>
<db 1>

-t strips the trailing newline from each element (you almost always want it). readarray is the same builtin under another name. Use it when you need random access or a count; use while read when the input is large or streaming.

read from the terminal

read -r -p "Deploy to prod? [y/N] " answer
[[ $answer == [yY] ]] || exit 1

read -r -s -p "Password: " pw; echo

-p prints a prompt (to stderr), -s does not show what you type, -t 10 gives up after ten seconds. In a script that might run without a terminal (from cron or a systemd timer), check first: [[ -t 0 ]] is true only when stdin (fd 0) is a terminal.

while read, or something faster?

A shell loop runs its body once per line, in the shell - fine for tens of thousands of lines, slow for millions, and every external program in the body (date, grep) starts a new process per line. When the job is "compute over columns", a text tool like awk (Ch 7) is one process and far faster. Reach for while read when each line drives an action: a command, a file operation.

What you can now do

Why it helps

Reading files and command output line by line is in almost every operational script: a list of hosts, a CSV of users to create, the output of another command, a config file. The bugs are subtle and ship: backslashes eaten in Windows paths, indentation stripped from config lines, the last line skipped because the file had no trailing newline, and counters or settings lost because the loop ran in a pipe subshell.

The "cat conf | while read loses every setting" bug is a real incident pattern. Knowing < <(cmd), find -print0 with read -d '', and when a per-line shell loop is simply too slow makes your scripts correct with odd input and usable with large input.

Commands in this lesson

printf echo mapfile readarray

FAQ

Why is IFS= placed before read and not on its own line?

A variable assignment directly before a command applies only to that command's environment. IFS= read -r line sets IFS to empty just for that read, so leading and trailing whitespace is kept, while the rest of the script keeps its normal IFS. Setting IFS= on its own line would change word splitting for everything after it, a common source of confusing bugs.

Why does my while loop skip the last line of a file?

read returns non-zero when it hits end of file, even if it read a partial line into the variable. A last line without a trailing newline therefore ends the loop before the body runs. Use while IFS= read -r line || [[ -n $line ]]; do to process that fragment. Files produced by some editors, Windows tools or printf without \n commonly lack the final newline.

Why are my variables empty after a while loop that reads from a pipe?

Each part of a pipeline runs in a subshell, so cmd | while read ...; do count=...; done sets count in a copy of the shell that is discarded when the pipeline ends. Redirect instead: while ...; done < file for files, or done < <(cmd) using process substitution for command output. In scripts, shopt -s lastpipe also runs the last pipeline stage in the current shell.

When should I use mapfile instead of while read?

mapfile -t arr < file (or readarray) loads all lines into an array in one fast builtin call, useful when you need the count, random access, or to iterate several times. while read streams line by line with constant memory, better for large or unbounded input and when each line triggers an action. -t removes the trailing newline from each element, which you almost always want.

How do I safely loop over files found by find?

Use NUL separators, because filenames can contain spaces and even newlines, but never NUL: while IFS= read -r -d '' f; do ...; done < <(find /path -type f -print0). -print0 terminates each name with NUL and read -d '' reads up to NUL. Alternatively, let find run the action directly with -exec cmd {} +, which avoids the loop entirely.

In an interview Junior

What is the correct way to read a file line by line in bash, and why is each part there?

while IFS= read -r line; do
  printf '%s\n' "$line"
done < input.txt

Add || [[ -n $line ]] to the condition so a last line without a newline is not skipped. To split fields: while IFS=: read -r user _ uid _ ; do ...; done < /etc/passwd.

Also asked: A script reads its settings with cat config | while read k v, but afterwards every setting is empty. Why? · How do you split a line into fields in bash? · How do you loop safely over filenames that may contain spaces?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.