Why this chapter
Something is wrong, and the evidence is text: a log with 300,000 lines, a config file, a blob of JSON. You cannot read it all. You need to ask it questions - "how many requests failed?", "who sent them?", "what happened just before?" - and get the answer in one line. This chapter is the handful of tools that do that: grep to find lines, sort/uniq to count, sed to edit, awk to compute over columns, jq for JSON.
What you need to know already: pipes and redirection (1.7); sudo (1.11); quoting and if (6.3, 6.6); globs (6.6). You have used grep in passing before - now properly.
The data you will work with
nginx is a web server: a program that answers requests from browsers and other programs over HTTP (the language of the web). This box runs it in front of the orders and payments apps you met in Ch 3. It writes two logs:
/var/log/nginx/access.log- one line per request. A line looks like:
10.0.9.41 - - [22/Sep/2026:17:53:00 +0000] "GET /static/app.js HTTP/1.1" 200 3810 "-" "python-requests/2.32.3" 0.080
From left: the client IP address (who asked; IP addresses are the numeric addresses machines have on a network, 1.13), two unused -, the time, the request (method GET or POST, and the path asked for), the status code, the bytes sent back, the page it came from, the user agent (the program that asked), and the seconds it took.
/var/log/nginx/error.log- nginx's own problems, readable by root only.
Paths starting with /api/ belong to the apps' API: the set of URLs other programs call to use them (POST /api/payments/refund = "please refund this"). Each such path is an endpoint.
A status code is the three-digit result of a request: 2xx = OK, 3xx = "look elsewhere", 4xx = the client asked wrong (404 not found), 5xx = the server failed (500 = the app crashed; 502 and 504 = nginx got no usable answer, or none in time, from the app behind it). Chapter 0 counted 5xx as the "errors" signal; here you find out who and what.
grep: print the lines that match
grep PATTERN FILE prints every line of FILE that contains PATTERN. The pattern is a regular expression (regex, 6.10): mostly plain text, plus a few special characters (below, and fully in 7.3).
The flags, in the order you will need them
-c count matching LINES (not matches - a line with three hits counts once)
-i case-insensitive
-v invert: lines that do NOT match
-n prefix with line numbers
-l just the filenames that matched
-r recurse into directories
-w whole words only
-q quiet: exit 0 on first match, print nothing. For if-tests.
-E extended regex: + ? | ( ) without backslashes
-F fixed string - no regex at all. Faster, and safe for user input.
-o print only the matching part, not the whole line
-A N -B N -C N N lines after / before / around each match
The ones that earn their keep
grep -c ' 502 ' access.log
-c: print only how many lines matched. The spaces around 502 make it match the status field and not, say, a byte count of 5021. Better than grep ' 502 ' access.log | wc -l - one process instead of two, and it is what -c is for. (shellcheck calls cat file | grep a "useless cat", SC2002.)
grep -B2 -A2 'upstream timed out' error.log
-B2 = also print 2 lines Before each match, -A2 = 2 lines After. Context is everything in a log. The matching line tells you what; the two lines either side usually tell you why. This is the single most useful grep command in an incident. (In nginx's error log, the upstream is the program nginx passes the request on to - here the payments app. "upstream timed out" means that app did not answer in time.)
grep -oE '^[0-9]+\.[0-9]+\.[0-9]+\.[0-9]+' access.log
-o turns grep into an extractor: print only the matching part, not the whole line. -E lets you use + without a backslash. The pattern reads: at the start of the line (^), digits ([0-9]+), a real dot (\.), digits... - an IP address. Feed that into sort | uniq -c (7.4) and you have a count per IP.
if grep -q 'ERROR' app.log; then echo "errors found"; fi
-q for tests - no output, just an exit status (6.5).
Regex, briefly
. any character ^ start of line
* 0 or more $ end of line
+ 1 or more (-E) [] a character class, [^] negated
? 0 or 1 (-E) \. a literal dot
| alternation (-E) () grouping (-E)
{2,5} a count range (-E)
Without -E you need backslashes for + ? | ( ) { }, which is why grep -E is the default worth adopting. The next lesson (7.3) explains the two dialects.
Always single-quote the pattern (6.6). Unquoted, the shell expands * and $ before grep ever sees them.
What you can now do
- Count, filter and invert matches with
-c,-i,-v. - Show the lines around a match with
-B/-A, and pull out just the match with-o.