OnCallReady

Lesson 7.6 · Text Processing & jq · 11 min read

sed

In plain words

Imagine giving an assistant a stack of pages and an instruction card: "on every line, cross out read.timeout.ms=0 and write read.timeout.ms=5000", or "tear out every line that starts with #", or "only give me pages 3 to 6". The assistant goes through the stack once, top to bottom, applying the card to each line, and hands you the result. By default they hand you a copy; with -i they change the original, and with -i.bak they keep a photocopy of the original first.

sed is that stream editor. s/old/new/g substitutes, /pattern/d deletes lines, -n '3,6p' prints a range. In the lab you fix /etc/orders/app.conf with sudo sed -i.bak.

Why this lesson

A config value is wrong on twelve lines, or you need "lines 40 to 60 of this file", or the same file without its comments. Opening an editor every time does not scale - and does not work inside a script. sed (stream editor) reads text line by line, changes it by rules you give it, and prints the result.

What you need to know already: regex basics (7.1, 7.3); sudo for files under /etc (1.11); the orders app's app.conf from 3.13.

How sed runs

sed 'SCRIPT' FILE reads FILE one line at a time, applies SCRIPT to each line, and prints every line (changed or not) to stdout. The file itself is not changed unless you ask with -i (below). The script is always in single quotes.

Substitution

s/PATTERN/REPLACEMENT/FLAGS replaces text matching the regex PATTERN:

sed 's/old/new/'          the FIRST match on each line
sed 's/old/new/g'         every match
sed 's/old/new/gi'        ...case-insensitively
sed -E 's/([0-9]+)/[\1]/g'   extended regex, with a capture group

g = global (all matches, not just the first); i = ignore case.

& in the replacement is the whole match, and \1..\9 are capture groups (the text each ( ) in the pattern matched, as in 6.10):

sed 's/^/# /'                    comment out every line (^ = start of line)
sed -E 's|^(/swap.img)|#\1|'     comment out one specific line

The delimiter does not have to be /. s|a|b|, s#a#b#, s,a,b, all work, and using something other than / when the pattern contains paths saves you a forest of backslashes.

In place, with a backup

-i (in place) writes the result back into the file instead of printing it:

sed -i 's/old/new/g' file          edit in place, no backup
sed -i.bak 's/old/new/g' file      keeps the original as file.bak

Use -i.bak on anything under /etc. It costs nothing and it is the difference between "undo" and "restore from backup".

Two gotchas:

Addresses: which lines

An address in front of a command limits it to certain lines: a line number, a range 3,6, a regex /BEGIN/, or $ for the last line. The commands here are p (print), d (delete, i.e. do not print) and q (quit):

sed -n '3,6p' file            print lines 3-6 (-n suppresses the default print)
sed -n '/BEGIN/,/END/p' file  from one pattern to another
sed '1d' file                 delete the first line (a header)
sed '$d' file                 delete the last
sed '/^#/d; /^$/d' file       strip comments and blank lines
sed '/pattern/s/a/b/'         substitute only on matching lines
sed '10q' file                quit after line 10 (like head, but stops reading)

-n plus p is the "print only what I ask for" mode, and is most of what you use sed for beyond substitution. ; separates two commands in one script.

When not to use sed

For JSON use jq (7.11). For a column of a structured log, awk (7.8) is clearer. sed is for line-oriented text substitution - and the moment your sed expression needs three layers of backslashes, you wanted a different tool.

What you can now do

Why it helps

sed is how config changes happen without an editor: in a setup script for a new server, in a release script that stamps a version number into a file, in the runbook step "set the timeout to 5000 on every server". You will read these in teammates' scripts, and the bugs are predictable: a pattern with an unescaped dot that also changes the wrong line, a / in a path that breaks the expression, sed -i with no backup on a file in /etc. On an incident box, sed -n '/10:14/,/10:20/p' pulls out a time window from a log in one line. Knowing that sed -i rewrites the file via a temp file in the same directory also explains the permission error you will hit on a file you own in a directory you do not.

Commands in this lesson

echo printf

FAQ

Why does sed only replace the first match on each line?

That is the default of s///: one replacement per line, the first match. Add the g flag (s/old/new/g) to replace every match on the line. You can also give a number, s/old/new/2 for only the second occurrence. It works per line: s/a/b/ without g still changes the first match on every line of the file, not just the first line.

How do I replace a path without escaping every slash?

Pick a different delimiter. The character right after s is the delimiter, so s|/usr/local/bin|/opt/bin|, s#a#b# and s,a,b, all work. Choose one that does not appear in the pattern or the replacement. Address patterns can use a custom delimiter too, but with a leading backslash: \|/var/log|d.

Why does sed -i fail with "couldn't open temporary file" on a file I own?

sed -i does not modify the file in place byte by byte. It writes the result to a new temporary file in the same directory and renames it over the original. Creating that temp file needs write permission on the directory. So a file you can write, in a directory you cannot, fails. It also means the inode changes: hard links and some file watchers do not see the edit the way you might expect.

Is sed -i the same on my Mac?

No. GNU sed on Ubuntu takes an optional suffix attached to the flag: -i or -i.bak. BSD sed on macOS requires the suffix as a separate argument, so sed -i '' 's/a/b/' f for no backup. A GNU-style sed -i 's/a/b/' f on a Mac treats the expression as the suffix and fails confusingly. -i.bak happens to work on both, another reason to always use it.

What does -n do, and why is it paired with p?

By default sed prints every line after processing it. -n turns that automatic print off, so only lines you explicitly print with p appear. sed -n '3,6p' prints lines 3 to 6; sed -n '/ERROR/p' acts like grep; sed -n 's/x/y/p' prints only lines where a substitution happened. Without -n, p prints matching lines twice.

In an interview Junior

How do you replace a string in a file in place, safely?

sed -i.bak 's/old/new/g' file: -i writes the result back into the file, .bak keeps the original as file.bak, and g replaces every match on a line, not just the first.

"Safely" means three things:

  1. Run it without -i first and read the output - sed has no undo.
  2. Keep a backup: -i.bak on anything under /etc, e.g. sudo sed -i.bak 's/read.timeout.ms=0/read.timeout.ms=5000/' /etc/orders/app.conf.
  3. Make the pattern specific: anchor it (^), escape dots, so it cannot hit a comment that mentions the same key.

If the pattern contains slashes, use another delimiter: s|/old/path|/new/path|. And sed -i needs write permission on the directory, because it writes a temporary file and renames it.

Also asked: How do you print only lines 40 to 60 of a file? · How do you strip comments and blank lines from a config file? · When would you use something other than sed?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.