Why this lesson
A pattern that works in grep -E silently matches nothing in sed. An "exact" IP search also matches a different IP. A pattern copied from the web does something odd. Regex bugs are quiet - you just get the wrong count. This lesson is the rules that make patterns behave the same everywhere.
What you need to know already: the regex basics from 6.10 and 7.1 (^ $ . + [0-9] ( | )); single quotes (6.6); grep -o and -c (7.1).
Three dialects, one idea
A literal character matches itself (a matches "a"). An operator (also called a metacharacter) has a special meaning (+ means "one or more"). Every tool in this chapter speaks regular expressions, but not the same dialect - they disagree about which characters are operators. (Perl, in the table, is a programming language whose regex style many others copied. Escaping means putting a backslash in front of a character to flip its meaning: \. is a literal dot, \( in BRE starts a group.)
BRE basic grep, sed (default) + ? | ( ) { } are LITERAL unless escaped
ERE extended grep -E, sed -E, awk + ? | ( ) { } are operators
PCRE Perl grep -P adds \d \s \w, lazy *?, lookarounds
The same pattern means different things in each:
$ echo 'foo(bar)' | grep 'o(b' # BRE: ( is a literal paren
foo(bar)
$ echo 'foo(bar)' | grep -E 'o\(b' # ERE: ( is grouping, so escape it
foo(bar)
$ echo aaa | grep -oE 'a{2,}' # ERE: {2,} is a count
aaa
$ echo aaa | grep -o 'a\{2,\}' # BRE: the same count needs backslashes
aaa
Use -E everywhere (grep -E, sed -E); awk is always ERE. Then + ? | ( ) { } mean what you expect and the patterns read the same across tools.
The pieces
. any one character \. a literal dot
* 0 or more of the previous + 1 or more (ERE)
? 0 or 1 (ERE) {n,m} between n and m (ERE)
^ $ start / end of line \b word boundary (GNU, and -P)
[abc] one of a, b, c [^abc] anything but
[a-z0-9] ranges a|b alternation (ERE)
( ) group, and capture (ERE) \1 back-reference to group 1
A character class [...] matches exactly one character from the set inside. A quantifier (* + ? {n,m}) says how many times the thing just before it may repeat. An anchor (^ $) matches a position, not a character.
POSIX character classes are named sets you put inside a bracket. They work in all three dialects and do not depend on the language settings (the locale) the way ranges like [a-z] can:
[[:digit:]] [[:alpha:]] [[:alnum:]] [[:space:]] [[:upper:]] [[:punct:]]
Note the double brackets: [:digit:] is only meaningful inside a bracket expression. grep '[:digit:]' matches the characters :, d, i, ... - GNU grep even warns you: grep: character class syntax is [[:space:]], not [:space:].
The dot is the classic bug
(sed 's/A/B/g' replaces every match of A with B; sed gets its own lesson in 7.6.)
$ echo "a.b.c" | sed 's/./X/g'
XXXXX
$ echo "a.b.c" | sed 's/\./X/g'
aXbXc
. matches any character, so an IP pattern 10.0.4.17 also matches 10a0b4c17 and 110.0.4.170. For exact matches: escape the dots and anchor it, or skip regex entirely:
grep -E '^10\.0\.4\.17 ' access.log anchored, dots escaped, space after
grep -F '10.0.4.17' access.log -F: fixed string, no regex at all
awk '$1 == "10.0.4.17"' access.log compare a field exactly - often best
Matching whole words and whole fields
$ cd /var/log/nginx
$ grep -c 'GET' access.log # would also match "TARGET", "GETTING"
347
$ grep -wc 'GET' access.log # -w: GET as a whole word
347
Same count here, because this log has no "TARGET" in it - but on a different file the first one quietly over-counts.
For logs, the field is usually the real unit. grep ' 502 ' (spaces around it) avoids a byte count of 50213, but awk '$9 == 502' says exactly what you mean.
Greedy
* and + are greedy: they match as much as possible:
$ echo '"GET /a" 200 "Mozilla/5.0"' | grep -oE '".*"'
"GET /a" 200 "Mozilla/5.0"
$ echo '"GET /a" 200 "Mozilla/5.0"' | grep -oE '"[^"]*"'
"GET /a"
"Mozilla/5.0"
"[^"]*" - a quote, anything that is not a quote, a quote - is the standard way to match "one quoted field" ([^"] is a negated class: any character except "). PCRE's lazy ".*?" (match as little as possible) also works with grep -P, but the negated class works everywhere.
Capture groups and back-references
$ echo "2026-09-22 key=val" | sed -E 's/([0-9]{4})-([0-9]{2})-([0-9]{2})/\3.\2.\1/'
22.09.2026 key=val
$ echo "a1b22c333" | grep -oE '[0-9]+'
1
22
333
( ) captures, \1..\9 in the replacement insert what was captured, & is the whole match (sed, 7.6). A back-reference is the same \1 used inside the pattern itself: \(ab\)\1 matches "abab". grep -o prints each match on its own line - the extraction half of every counting pipeline.
PCRE, when you need it
grep -P '\d{3} \d+' \d is PCRE; in ERE write [0-9]
grep -oP '(?<=user=)\w+' lookbehind: the word after user=
grep -oP 'took \K\d+(?=ms)' \K drops what came before
A lookaround checks what is before ((?<=...), lookbehind) or after ((?=...), lookahead) the match without including it. \w is a word character (letter, digit, _). Lookarounds are the reason to reach for -P: "print the value after this key" without printing the key. \d in grep -E does not mean digit (it is a literal d in POSIX regex) - a quiet bug when copying patterns from other languages.
Quote the pattern
grep -E '^[0-9]+$' file # single quotes: the shell leaves it alone
grep -E ^[0-9]+$ file # the shell may glob [0-9]+ against filenames
Always single quotes, unless you are deliberately inserting a shell variable - and then only if the variable holds a regex you trust. For text someone else typed, grep -F (fixed string) or awk -v x="$val" '$1 == x' (-v passes a shell value into awk as a variable) treats it as plain text, so a stray . or * in it cannot change the pattern.
What you can now do
- Write a pattern once with
-Eand use it in grep, sed and awk. - Match an IP or a field exactly: escaped dots and anchors,
-F, or$1 == "x". - Extract quoted fields with
"[^"]*", and reach for-Ponly for lookarounds.