OnCallReady

Lesson 7.3 · Text Processing & jq · 21 min read

Regular expressions: BRE, ERE and PCRE

In plain words

Think of a regex as a stencil for text. A plain stencil only matches one exact word. A clever stencil has holes that say "any letter here", "one or more digits here", "this must be the start of the line". Hold the stencil over each line; if it fits, the line matches.

The catch is that there are three workshops making stencils with slightly different symbols. In basic regex (BRE, plain grep and sed) a + is a plain plus sign unless you add a backslash. In extended regex (ERE, grep -E, sed -E, awk) + means "one or more". Perl regex (PCRE, grep -P) adds shortcuts like \d and lookarounds. This lesson is about knowing which workshop you are in.

Why this lesson

A pattern that works in grep -E silently matches nothing in sed. An "exact" IP search also matches a different IP. A pattern copied from the web does something odd. Regex bugs are quiet - you just get the wrong count. This lesson is the rules that make patterns behave the same everywhere.

What you need to know already: the regex basics from 6.10 and 7.1 (^ $ . + [0-9] ( | )); single quotes (6.6); grep -o and -c (7.1).

Three dialects, one idea

A literal character matches itself (a matches "a"). An operator (also called a metacharacter) has a special meaning (+ means "one or more"). Every tool in this chapter speaks regular expressions, but not the same dialect - they disagree about which characters are operators. (Perl, in the table, is a programming language whose regex style many others copied. Escaping means putting a backslash in front of a character to flip its meaning: \. is a literal dot, \( in BRE starts a group.)

BRE   basic      grep, sed (default)       + ? | ( ) { } are LITERAL unless escaped
ERE   extended   grep -E, sed -E, awk      + ? | ( ) { } are operators
PCRE  Perl       grep -P                   adds \d \s \w, lazy *?, lookarounds

The same pattern means different things in each:

$ echo 'foo(bar)' | grep 'o(b'          # BRE: ( is a literal paren
foo(bar)
$ echo 'foo(bar)' | grep -E 'o\(b'      # ERE: ( is grouping, so escape it
foo(bar)
$ echo aaa | grep -oE 'a{2,}'            # ERE: {2,} is a count
aaa
$ echo aaa | grep -o 'a\{2,\}'           # BRE: the same count needs backslashes
aaa

Use -E everywhere (grep -E, sed -E); awk is always ERE. Then + ? | ( ) { } mean what you expect and the patterns read the same across tools.

The pieces

.          any one character              \.       a literal dot
*          0 or more of the previous      +        1 or more (ERE)
?          0 or 1 (ERE)                   {n,m}    between n and m (ERE)
^  $       start / end of line            \b       word boundary (GNU, and -P)
[abc]      one of a, b, c                 [^abc]   anything but
[a-z0-9]   ranges                         a|b      alternation (ERE)
( )        group, and capture (ERE)       \1       back-reference to group 1

A character class [...] matches exactly one character from the set inside. A quantifier (* + ? {n,m}) says how many times the thing just before it may repeat. An anchor (^ $) matches a position, not a character.

POSIX character classes are named sets you put inside a bracket. They work in all three dialects and do not depend on the language settings (the locale) the way ranges like [a-z] can:

[[:digit:]]  [[:alpha:]]  [[:alnum:]]  [[:space:]]  [[:upper:]]  [[:punct:]]

Note the double brackets: [:digit:] is only meaningful inside a bracket expression. grep '[:digit:]' matches the characters :, d, i, ... - GNU grep even warns you: grep: character class syntax is [[:space:]], not [:space:].

The dot is the classic bug

(sed 's/A/B/g' replaces every match of A with B; sed gets its own lesson in 7.6.)

$ echo "a.b.c" | sed 's/./X/g'
XXXXX
$ echo "a.b.c" | sed 's/\./X/g'
aXbXc

. matches any character, so an IP pattern 10.0.4.17 also matches 10a0b4c17 and 110.0.4.170. For exact matches: escape the dots and anchor it, or skip regex entirely:

grep -E '^10\.0\.4\.17 ' access.log       anchored, dots escaped, space after
grep -F '10.0.4.17' access.log            -F: fixed string, no regex at all
awk '$1 == "10.0.4.17"' access.log        compare a field exactly - often best

Matching whole words and whole fields

$ cd /var/log/nginx
$ grep -c 'GET' access.log       # would also match "TARGET", "GETTING"
347
$ grep -wc 'GET' access.log      # -w: GET as a whole word
347

Same count here, because this log has no "TARGET" in it - but on a different file the first one quietly over-counts.

For logs, the field is usually the real unit. grep ' 502 ' (spaces around it) avoids a byte count of 50213, but awk '$9 == 502' says exactly what you mean.

Greedy

* and + are greedy: they match as much as possible:

$ echo '"GET /a" 200 "Mozilla/5.0"' | grep -oE '".*"'
"GET /a" 200 "Mozilla/5.0"
$ echo '"GET /a" 200 "Mozilla/5.0"' | grep -oE '"[^"]*"'
"GET /a"
"Mozilla/5.0"

"[^"]*" - a quote, anything that is not a quote, a quote - is the standard way to match "one quoted field" ([^"] is a negated class: any character except "). PCRE's lazy ".*?" (match as little as possible) also works with grep -P, but the negated class works everywhere.

Capture groups and back-references

$ echo "2026-09-22 key=val" | sed -E 's/([0-9]{4})-([0-9]{2})-([0-9]{2})/\3.\2.\1/'
22.09.2026 key=val
$ echo "a1b22c333" | grep -oE '[0-9]+'
1
22
333

( ) captures, \1..\9 in the replacement insert what was captured, & is the whole match (sed, 7.6). A back-reference is the same \1 used inside the pattern itself: \(ab\)\1 matches "abab". grep -o prints each match on its own line - the extraction half of every counting pipeline.

PCRE, when you need it

grep -P '\d{3} \d+'                    \d is PCRE; in ERE write [0-9]
grep -oP '(?<=user=)\w+'               lookbehind: the word after user=
grep -oP 'took \K\d+(?=ms)'             \K drops what came before

A lookaround checks what is before ((?<=...), lookbehind) or after ((?=...), lookahead) the match without including it. \w is a word character (letter, digit, _). Lookarounds are the reason to reach for -P: "print the value after this key" without printing the key. \d in grep -E does not mean digit (it is a literal d in POSIX regex) - a quiet bug when copying patterns from other languages.

Quote the pattern

grep -E '^[0-9]+$' file      # single quotes: the shell leaves it alone
grep -E ^[0-9]+$ file        # the shell may glob [0-9]+ against filenames

Always single quotes, unless you are deliberately inserting a shell variable - and then only if the variable holds a regex you trust. For text someone else typed, grep -F (fixed string) or awk -v x="$val" '$1 == x' (-v passes a shell value into awk as a variable) treats it as plain text, so a stray . or * in it cannot change the pattern.

What you can now do

Why it helps

Regex shows up everywhere on a platform team, not only in grep: bash's [[ =~ ]], sed and awk, nginx location blocks, input validation in scripts, and almost every alerting and log-search tool you will meet later. The bugs are always the same few. An unescaped dot makes 10.0.4.17 also match 110.0.4.170, and your count includes the wrong host. A greedy .* swallows half a log line. A \d copied from a Python snippet silently means the letter d in grep -E. Knowing the three dialects lets you read a teammate's pattern in a review and see that it matches more (or less) than they think, before it causes a wrong answer at 3am.

Commands in this lesson

echo cd grep

FAQ

Why does \d not work in grep -E?

Because \d is not part of POSIX regex. It comes from Perl and was adopted by most programming languages and PCRE. In grep -E it is treated as a literal d (recent GNU grep also prints a "stray \ before d" warning). Write [0-9] or [[:digit:]] in ERE, or switch to grep -P when you really want Perl syntax. awk and sed have no \d either.

What is the difference between [:digit:] and [[:digit:]]?

[:digit:] is only meaningful inside a bracket expression, so the full syntax is [[:digit:]]: the outer brackets are the character class, the inner [:digit:] is the named set. Written alone, [:digit:] is a bracket expression containing the characters : d i g t, which is almost never what you meant. GNU grep catches this and warns you with the correct syntax.

Why does .* match far more than I expected?

* and + are greedy: they match as much as possible while still letting the whole pattern succeed. ".*" on a line with several quoted fields matches from the first quote to the last one. The portable fix is a negated class, "[^"]*", which cannot cross a quote. PCRE's lazy .*? also works, but only in grep -P and languages with Perl-style regex, not in sed or awk.

When should I just use grep -F or awk instead of a regex?

Whenever you want an exact value. grep -F '10.0.4.17' treats the dots literally, and awk '$1 == "10.0.4.17"' compares one field exactly, so it cannot accidentally match a substring elsewhere on the line. Regex is for patterns, like "any 5xx" or "anything after user=". Also use -F for user-supplied strings: it removes any chance of the input being interpreted as regex.

What are capture groups and back-references for?

Parentheses capture what they matched, and \1 to \9 refer back to it. In sed replacements that lets you rearrange text: turn 2026-09-22 into 22.09.2026, or wrap a value in quotes while keeping it. & in the replacement is the whole match. In ERE (sed -E) you write ( ); in BRE you would need \( \). grep itself does not replace, so groups matter most in sed and in jq's sub/capture.

In an interview Junior

What is the difference between basic and extended regular expressions?

The operators are the same; which characters need a backslash is reversed.

., *, ^, $ and [...] mean the same in all of them. I use -E everywhere so a pattern reads the same in grep, sed and awk. Two classic bugs: \d in grep -E is a literal d, not a digit (write [0-9]); and an unescaped . matches any character, so 10.0.4.17 also matches 110.0.4.170 - escape the dots and anchor it, or use grep -F.

Also asked: Write a regex for an IPv4 address in a log line. What are the pitfalls? · What does "greedy" mean, and how do you match one quoted field? · What are capture groups, and how do you use them in sed?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.