OnCallReady

Lesson 7.4 · Text Processing & jq · 9 min read

cut, sort, uniq, tr, column

In plain words

Think of a sorting machine at a post office. One machine reads only the postcode off each letter (cut). The next puts letters with the same postcode next to each other (sort). The next counts each run of identical postcodes (uniq -c). The last one orders the piles from biggest to smallest (sort -rn), and you look at the top ten (head). Each machine does one small job and passes everything on.

On the command line the conveyor belt is the pipe |, and sort | uniq -c | sort -rn | head is the most reused four-machine line there is. tr cleans characters on the way, column -t lines the output up so a human can read it.

Why this lesson

"Which IP sent the most requests?" "Which path fails most?" These are all the same question - what is the most common X? - and a chain of four tiny tools answers every one of them. Each tool does one small job on a stream of lines (text flowing from one command to the next through pipes) and passes the result on.

What you need to know already: pipes (1.7); fields and delimiters (6.8); the access log's layout (7.1).

The small tools between the big ones

The flags you will use:

cut -d: -f1 /etc/passwd          field 1, colon-delimited
cut -d: -f1,3 /etc/passwd        fields 1 and 3
cut -c1-10                       characters instead of fields

sort            lexical (as text, character by character)
sort -n         numeric
sort -h         human sizes: 2K < 1M < 1G  (pairs with du -h)
sort -r         reverse
sort -u         unique (= sort | uniq)
sort -k2 -rn    on field 2, numeric, descending
sort -t: -k3 -n on field 3, colon-separated

uniq -c         prefix each line with its count
uniq -d         only duplicated lines
uniq -u         only unique lines

tr -d '\r'      delete characters (\r ends every line in Windows files)
tr -s ' '       squeeze runs of spaces into one
tr 'A-Z' 'a-z'  lowercase

column -t       align into columns
wc -l           count lines

**uniq only collapses adjacent duplicates.** It is a streaming tool with no memory. So it is always sort | uniq -c, never uniq -c alone.

cut cannot treat runs of spaces as one delimiter. With a log that is aligned with variable spacing, cut -d' ' -f3 gives you an empty field. Either tr -s ' ' first, or - better - use awk (7.8), which splits on runs of spaces by default.

The idiom

... | sort | uniq -c | sort -rn | head -10

Extract the thing, sort so duplicates are adjacent, uniq -c to count each run, sort -rn by that count descending, head -10 to take the top ten. That four-stage pipeline answers "what is the most common X" for every X you will ever be asked about: IPs, status codes, paths, error messages, user agents.

The extraction step is usually one tiny awk program. You only need this much of awk now (7.8 explains it properly): awk '{print $1}' prints the first field of every line, splitting on spaces. $7 would be the seventh field.

awk '{print $1}' access.log | sort | uniq -c | sort -rn | head -10

That is the classic interview question - top 10 IPs in one line - and it generalises by changing one field number. In the access log: $1 the client, $7 the path, $9 the status.

What you can now do

Why it helps

"What is the most common X" is the question behind half of all log investigations: which IP, which status code, which URL, which error message, which user agent. The pipeline in this lesson answers all of them by changing one field number, in seconds, on any box, without a logging platform. You will also use sort -h with du -h every time a disk fills up, cut -d: -f1 /etc/passwd when auditing users, and tr -d '\r' the first time someone commits a script or CSV edited on Windows and it breaks with a mysterious $'\r': command not found. In interviews the top-10 one-liner is practically guaranteed.

Commands in this lesson

printf

FAQ

Why do I need sort before uniq?

uniq only compares each line with the one directly before it. It keeps no memory of earlier lines, which is why it can stream through huge inputs. So a b a stays three lines. sort puts identical lines next to each other first, then uniq can collapse them. If you only need the distinct values and not the counts, sort -u does both in one step.

Why does cut give me an empty field?

cut -d' ' treats every single space as a separator, so two spaces in a row create an empty field between them. Aligned output like ps, df or padded logs has runs of spaces, so the field numbers shift. Either squeeze the spaces first with tr -s ' ', or use awk '{print $3}', which splits on runs of whitespace by default. For whitespace-aligned text awk is almost always the right tool.

What is the difference between sort -n and sort -h?

-n compares the leading number numerically, so 10 comes after 9 instead of before it (plain lexical sort gives 1, 10, 2). -h understands human-readable size suffixes, so 900K sorts below 2M and 2M below 1G. Use -h with anything from du -h, df -h or ls -lh; -n would treat 1G as 1 and put it below 900K.

How do I sort by a column other than the first?

-k picks the key and -t the separator: sort -t: -k3 -n /etc/passwd sorts by UID. Note that -k3 means "from field 3 to the end of the line"; -k3,3 limits it to exactly field 3, which matters when later fields could break ties unexpectedly. Flags can be attached per key: sort -k2,2nr -k1,1 sorts by field 2 numerically descending, then by field 1.

What does tr actually do, and why can it not replace words?

tr translates or deletes single characters, one for one. tr 'a-z' 'A-Z' maps each lowercase letter to its uppercase partner, tr -d '\r' deletes carriage returns, tr -s ' ' squeezes repeats. It has no idea of words or patterns, and it only reads stdin, never a filename. To replace a word or a pattern, use sed 's/old/new/g'.

In an interview Junior

How do you find the most common HTTP status codes in an access log?

awk '{print $9}' access.log | sort | uniq -c | sort -rn

In nginx's log the status code is the ninth space-separated field, so awk '{print $9}' extracts it. sort groups identical codes together - needed because uniq has no memory and only collapses adjacent duplicates. uniq -c counts each run. sort -rn orders by that count, numerically (-n) and biggest first (-r); without -n, "9" would sort above "10".

Check the field number against one real line first - a different log format shifts every field. And cut -d' ' -f9 is not a replacement for awk here: cut treats every single space as a delimiter, so runs of spaces give empty fields.

Also asked: What is the difference between sort | uniq and sort -u? · Why does uniq almost always come after sort? · When does cut fail where awk works?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.