OnCallReady

Lesson 7.1 · Text Processing & jq · 14 min read

grep

In plain words

Picture a librarian with a highlighter who reads a whole book line by line and hands you only the lines containing the word you asked for. You can ask for extras: "just tell me how many lines" (-c), "give me two lines before and after, so I understand the sentence" (-B2 -A2), "ignore capital letters" (-i), or "just nod if it is in there at all" (-q).

grep is that librarian for files. It reads text line by line, and every line that matches your pattern is printed. On oncall-lab you point it at the nginx access and error logs to count 502s, pull context around upstream timed out, and extract IPs with -oE.

Why this chapter

Something is wrong, and the evidence is text: a log with 300,000 lines, a config file, a blob of JSON. You cannot read it all. You need to ask it questions - "how many requests failed?", "who sent them?", "what happened just before?" - and get the answer in one line. This chapter is the handful of tools that do that: grep to find lines, sort/uniq to count, sed to edit, awk to compute over columns, jq for JSON.

What you need to know already: pipes and redirection (1.7); sudo (1.11); quoting and if (6.3, 6.6); globs (6.6). You have used grep in passing before - now properly.

The data you will work with

nginx is a web server: a program that answers requests from browsers and other programs over HTTP (the language of the web). This box runs it in front of the orders and payments apps you met in Ch 3. It writes two logs:

10.0.9.41 - - [22/Sep/2026:17:53:00 +0000] "GET /static/app.js HTTP/1.1" 200 3810 "-" "python-requests/2.32.3" 0.080

From left: the client IP address (who asked; IP addresses are the numeric addresses machines have on a network, 1.13), two unused -, the time, the request (method GET or POST, and the path asked for), the status code, the bytes sent back, the page it came from, the user agent (the program that asked), and the seconds it took.

Paths starting with /api/ belong to the apps' API: the set of URLs other programs call to use them (POST /api/payments/refund = "please refund this"). Each such path is an endpoint.

A status code is the three-digit result of a request: 2xx = OK, 3xx = "look elsewhere", 4xx = the client asked wrong (404 not found), 5xx = the server failed (500 = the app crashed; 502 and 504 = nginx got no usable answer, or none in time, from the app behind it). Chapter 0 counted 5xx as the "errors" signal; here you find out who and what.

grep: print the lines that match

grep PATTERN FILE prints every line of FILE that contains PATTERN. The pattern is a regular expression (regex, 6.10): mostly plain text, plus a few special characters (below, and fully in 7.3).

The flags, in the order you will need them

-c            count matching LINES (not matches - a line with three hits counts once)
-i            case-insensitive
-v            invert: lines that do NOT match
-n            prefix with line numbers
-l            just the filenames that matched
-r            recurse into directories
-w            whole words only
-q            quiet: exit 0 on first match, print nothing. For if-tests.
-E            extended regex: + ? | ( ) without backslashes
-F            fixed string - no regex at all. Faster, and safe for user input.
-o            print only the matching part, not the whole line
-A N -B N -C N   N lines after / before / around each match

The ones that earn their keep

grep -c ' 502 ' access.log

-c: print only how many lines matched. The spaces around 502 make it match the status field and not, say, a byte count of 5021. Better than grep ' 502 ' access.log | wc -l - one process instead of two, and it is what -c is for. (shellcheck calls cat file | grep a "useless cat", SC2002.)

grep -B2 -A2 'upstream timed out' error.log

-B2 = also print 2 lines Before each match, -A2 = 2 lines After. Context is everything in a log. The matching line tells you what; the two lines either side usually tell you why. This is the single most useful grep command in an incident. (In nginx's error log, the upstream is the program nginx passes the request on to - here the payments app. "upstream timed out" means that app did not answer in time.)

grep -oE '^[0-9]+\.[0-9]+\.[0-9]+\.[0-9]+' access.log

-o turns grep into an extractor: print only the matching part, not the whole line. -E lets you use + without a backslash. The pattern reads: at the start of the line (^), digits ([0-9]+), a real dot (\.), digits... - an IP address. Feed that into sort | uniq -c (7.4) and you have a count per IP.

if grep -q 'ERROR' app.log; then echo "errors found"; fi

-q for tests - no output, just an exit status (6.5).

Regex, briefly

.        any character         ^  start of line
*        0 or more             $  end of line
+        1 or more   (-E)      [] a character class, [^] negated
?        0 or 1      (-E)      \. a literal dot
|        alternation (-E)      () grouping (-E)
{2,5}    a count range (-E)

Without -E you need backslashes for + ? | ( ) { }, which is why grep -E is the default worth adopting. The next lesson (7.3) explains the two dialects.

Always single-quote the pattern (6.6). Unquoted, the shell expands * and $ before grep ever sees them.

What you can now do

Why it helps

During an incident, grep is usually the first tool you touch. The alert says the orders API is failing; grep -c ' 502 ' /var/log/nginx/access.log gives you the size of the problem in a second, and grep -B2 -A2 'upstream timed out' error.log shows which upstream and when. In scripts, grep -q is how you test "does this file contain X" without printing anything, for example checking that a generated config contains the setting you expect. grep -r across a repo is how you find every place a config key or env var is used before you rename it. And -F matters the moment you search for a literal string with dots or brackets in it, like an IP or a version number.

Commands in this lesson

grep

FAQ

Does grep -c count matches or lines?

Lines. A line with the pattern three times counts once. If you need the number of matches, use grep -o pattern file | wc -l, since -o prints each match on its own line. For log work lines are usually what you want, because one log line is one request or one event, so grep -c ' 502 ' is the number of requests that returned 502.

Why is grep -c better than grep | wc -l?

They give the same number, but grep -c is one process and says what you mean. The real anti-pattern is cat file | grep x, which shellcheck flags as SC2002 (useless use of cat): grep can read the file itself. It is not a correctness bug, but it adds a process and hides the filename from grep, so options like -l or -H cannot report it.

What is the difference between grep, grep -E and grep -F?

Plain grep uses basic regex (BRE), where + ? | ( ) { } are literal characters unless backslashed. grep -E uses extended regex, where they are operators, which is what most people expect. grep -F turns regex off completely and searches for the exact string, so a dot is just a dot. Use -F for IPs, version strings and anything a user typed; it is also faster on large files.

Why does my grep pattern behave strangely without quotes?

The shell sees the pattern first. Unquoted, *, ? and [...] are glob characters, and if they match filenames in the current directory the shell replaces the pattern with those names before grep runs. $ starts a variable expansion. Single quotes pass the pattern through untouched. Use double quotes only when you deliberately want a shell variable inside the pattern.

Why does grep find nothing in a file I know contains the text?

Common causes: the file is compressed (use zgrep for rotated .gz logs), it has Windows line endings so $ anchors fail, the case differs (-i), you have no permission and the error went to stderr (nginx's error.log needs sudo on oncall-lab), or the pattern has an unescaped special character like . or (. Try grep -F with a short literal piece first, then build the pattern back up.

In an interview Junior

How do you find an error in a big log and see what led up to it?

grep prints the lines that match a pattern; the flags do the rest:

Always single-quote the pattern, so the shell does not expand * or $ first. The error log is root-only: sudo grep.

Also asked: How do you search a whole directory and list only the files that contain a string? · How would you use grep inside an if in a script? · What is the difference between grep -c and grep ... | wc -l?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.