OnCallReady

Chapter 7 Text Processing & jq

grep, regular expressions, sed, awk, jq and the little tools between them - turning a log and a blob of JSON into an answer.

In plain words

Imagine a huge pile of receipts from a shop, thousands of them, and someone asks "who bought the most yesterday?". You would not read every one. You would pull out yesterday's receipts, write the customer name on a sticky note for each, sort the notes into piles and count the piles. Each step is a simple job done by one helper, and the helpers pass the pile down the line.

That is this chapter. The receipts are log lines and JSON documents. The helpers are grep (pull out the matching lines), cut and awk (copy one field), sort and uniq -c (pile and count), sed (rewrite text), jq (the same jobs for JSON) and xargs (hand the result to a command as arguments). The shell pipe | is the line of helpers.

Why it matters on call

Most of your first minutes in an incident are spent turning text into an answer. The payments API starts returning 502s: grep -c ' 502 ' tells you how many, grep -B2 -A2 on the nginx error log tells you why, and one awk | sort | uniq -c | sort -rn line tells you which client is doing it. jq is how you pull the entries that are not Running out of a JSON list of hundreds without scrolling. sed -i.bak is how you fix a config value on a box without opening an editor. In interviews, "top 10 IPs in an nginx log in one line" and "output a TSV from this JSON filtered by a field" are standard live exercises. And in teammates' scripts you will read these tools constantly, so knowing their gotchas is how you spot the bug in a review.

Lessons

  1. grep
  2. Regular expressions: BRE, ERE and PCRE
  3. cut, sort, uniq, tr, column
  4. sed
  5. awk
  6. jq
  7. jq beyond select
  8. xargs

10 hands-on labs (missions, incidents and drills) run in the terminal: Open this chapter in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.

Questions people ask

When do I use grep, sed or awk?

grep answers "which lines match?". sed answers "change this text on these lines" (substitution, deletion, printing a range). awk answers "treat each line as fields and compute something": compare a column numerically, sum a column, count by a key. A rough rule: if you need a column or arithmetic, it is awk; if you need a replacement, it is sed; if you only need to filter, it is grep. JSON is none of them: use jq.

Why not just write a Python script?

For a one-off question during an incident, the shell pipeline is faster to type than opening an editor, and it runs on every box: grep, sed, awk and sort are on every Linux server, while Python may not be installed on a minimal one. When the logic grows past a line or two, or you need proper error handling and tests, a script is the right call. The skill is knowing where that line is.

Are these the same tools as on my Mac?

No, and it bites. macOS ships BSD versions: sed -i on a Mac requires a suffix argument (sed -i '' 's/a/b/' f), BSD grep has no -P, and date, stat and xargs flags differ. Ubuntu has the GNU versions, which is what this chapter teaches and what your servers use. If you script on the Mac, either install the GNU tools via Homebrew (gsed, ggrep) or test on a Linux box.

Why does my pipeline print nothing and not fail?

Because an empty result is a valid result. grep exits 1 when nothing matches, but in a pipeline the shell reports only the last command's status unless you use set -o pipefail. Common causes: the pattern is wrong (unescaped dot, missing -E), the file has Windows line endings (\r at the end of every line), or you are reading a file you cannot read and the error went to stderr. Run each stage alone and add stages one at a time.

Do I really need to memorise regex?

A small core, yes: ., *, +, ?, ^, $, [...], [^...], |, ( ) and {n,m}. With those, plus knowing that -E turns on the modern syntax, you can read and write nearly every pattern you will meet in logs, configs and scripts. Lookarounds and \K from PCRE are worth recognising, not memorising. The same regex knowledge carries into nginx config, [[ =~ ]] in bash and most programming languages.