OnCallReady

Lesson 7.8 · Text Processing & jq · 15 min read

awk

In plain words

Imagine a spreadsheet where every line of a text file is a row and every word is a cell. awk lets you say: "for every row where column 9 starts with 5, add column 10 to a running total", and at the end, "print the total". Or: "keep a tally per value of column 1", like counting votes with tick marks next to each name.

That is awk: pattern { action } run on every line, with $1, $2, $NF as the cells, BEGIN and END for before and after, and associative arrays (c[$1]++) as the tally sheet. On oncall-lab you use it on /etc/passwd (users with UID >= 1000) and on the nginx access log (total bytes, average request time, counts per client and path).

Why this lesson

"Total bytes served", "average response time", "5xx count per client and path" - questions that need columns and arithmetic. grep only knows whole lines; cut cannot add. awk is a small programming language made for exactly this: it reads a file line by line, splits each line into fields, and lets you compare, add up and count them.

What you need to know already: fields and delimiters (6.8); regex (7.3); associative arrays (6.12) - awk has them too; printf formats (6.5); the access log's fields (7.1, 7.4).

awk is "for each line, if PATTERN then ACTION"

On your VM: Ubuntu's awk is mawk, not gawk (two versions of the same tool). Everything in this chapter works in both; gawk-only extras (PROCINFO["sorted_in"], strftime, gensub) need sudo apt install gawk.

awk 'PATTERN { ACTION }' file

The text in quotes is an awk program. For every line (awk calls it a record), awk checks the PATTERN; if it is true, it runs the ACTION in braces.

Leave out the pattern and the action runs on every line. Leave out the action and awk prints the matching line. So awk '/error/' works like grep, and awk '{print $2}' like a smarter cut.

The variables

awk fills these in for you on every line:

$0    the whole line          NF    number of fields on this line
$1..  the fields              NR    record (line) number, running total
$NF   the LAST field          FS    input field separator (default: whitespace)
$(NF-1)  second to last       OFS   output field separator

$1 here is awk's own "field 1", not a shell variable - that is why the awk program is always in single quotes, so the shell leaves $1 alone.

awk splits on runs of whitespace by default, which is precisely what cut cannot do. For any other delimiter, -F (field separator):

awk -F: '{print $1}' /etc/passwd
awk -F, '{print $3}' data.csv
awk 'BEGIN{FS=":"; OFS="\t"} {print $3, $1}' /etc/passwd

(BEGIN{...} runs once before the first line - below. OFS is what print a, b puts between a and b; \t is a tab.)

Patterns

~ means "matches this regex", !~ "does not match", == "is equal to":

awk '$3 >= 1000'                  numeric comparison on a field
awk '$9 ~ /^5/'                   field matches a regex
awk '$9 !~ /^2/'                  does not match
awk '/GET/ && $9 == 200'          combine
awk 'NR > 1'                      skip the header
awk 'NR==10, NR==20'              a range of lines

BEGIN, END and accumulation

BEGIN { ... } runs once before any line is read; END { ... } runs once after the last line. An accumulator is a variable you add to on every line and print in END. awk variables start at 0 (or empty) without being declared, and sum += $10 means "add field 10 to sum". This is where awk stops being "cut with extras":

awk '{sum += $10} END {print sum}' access.log            total bytes
awk '{s += $NF} END {printf "%.3f\n", s/NR}' access.log   average, 3 decimals
awk '{c[$1]++} END {for (ip in c) print c[ip], ip}' access.log

The second divides by NR, which in END is the number of lines read. %.3f is a printf format (6.5): a number with 3 decimals.

The last one is the important one. c[$1]++ uses an associative array (6.12

sees that value. for (ip in c) walks every key in END. That is sort | uniq -c in a single pass, and it can count anything - including combinations:

awk '$9 ~ /^5/ {c[$1" "$7]++} END {for (k in c) print c[k], k}' access.log \
  | sort -rn | head

"For every 5xx, count by client and path." $1" "$7 glues field 1, a space and field 7 into one string - a composite key. (The trailing \ continues the command on the next line.) Two greps and a cut could not do that.

printf in awk works like the shell's: printf "%-20s %6.2f\n", name, value - but with commas between the arguments.

awk versus the pipeline

sort | uniq -c | sort -rn is easier to type and easier to read. awk's array version is one pass instead of three, which matters on a large log, and lets you key on something computed rather than a whole field.

Use whichever is clearer until the log is big enough for it to matter.

What you can now do

Why it helps

awk is what you reach for when the question has a condition or arithmetic in it, which is most real questions: "how many 5xx per path", "average response time in the last hour", "which users have a UID above 1000", "sum the memory of all java processes from ps". It works on whitespace-aligned output from ps, df and free, where cut fails. You will see it in shell scripts and runbooks everywhere, often as awk '{print $2}' to grab a PID or a version, and you need to read those confidently. In an interview, being able to count by two fields at once in one awk line shows you can actually work with logs, not just grep them.

Commands in this lesson

echo awk

FAQ

What is the difference between NR and NF?

NR is the record number: how many lines awk has read so far, across all input files (FNR restarts per file). NF is the number of fields on the current line. So $NF is the last field, $(NF-1) the second to last, and NR > 1 skips a header line. END {print NR} is a line count, the same as wc -l.

Why use awk instead of cut?

awk splits on runs of whitespace by default, so aligned output with variable spacing (ps aux, df -h, free -m) just works, while cut -d' ' produces empty fields. awk can also compare fields ($3 >= 1000), do arithmetic, reorder fields, and filter and extract in one step. cut is fine for strictly delimited data like /etc/passwd, and is slightly simpler to read there.

Why is $3 >= 1000 comparing numbers and not text?

awk decides per comparison: if both sides look like numbers (a numeric constant, or a field whose content looks numeric), it compares numerically; otherwise as strings. $3 >= 1000 on /etc/passwd compares UIDs as numbers. That is also why the lesson's surprise happens: nobody has UID 65534, which is >= 1000, so a naive "real users" filter includes it. Add && $3 < 65534 or filter on the shell field.

Why is the order of for (k in c) random?

awk arrays are hash tables, and for (k in c) walks them in an unspecified order that depends on the implementation. If you need the output sorted, pipe it: ... END {for (k in c) print c[k], k}' | sort -rn. gawk has PROCINFO["sorted_in"] to control the order, but Ubuntu's default awk is mawk, which does not, so piping to sort is the portable answer.

Is awk a programming language?

Yes, a small one: variables, arrays, loops, if/else, functions, printf, string functions (sub, gsub, split, substr, length, match). That is why one-liners scale to real reports. The limit is readability: past a few statements, put the program in a file (awk -f report.awk) or switch to Python. Ubuntu installs mawk as awk; gawk (GNU awk) adds extras such as time functions.

In an interview Junior

How do you print one column of a command's output, and how would you add a column up?

awk '{print $2}' - awk reads line by line, splits each line into fields on runs of whitespace, and $2 is the second field ($1 the first, $NF the last). A pattern in front makes it a filter: awk '$9 ~ /^5/ {print $1}' access.log prints the client of every 5xx. For another delimiter, -F: awk -F: '{print $1}' /etc/passwd.

Adding up uses an accumulator and an END block, which runs once after the last line:

awk '{sum += $10} END {print sum}' access.log                total bytes
awk '{s += $NF} END {printf "%.3f\n", s/NR}' access.log      average time
awk '{c[$1]++} END {for (ip in c) print c[ip], ip}' access.log   count per IP

Variables start at 0 without being declared; NR in END is the number of lines read. Keep the program in single quotes so the shell leaves $1 alone.

Also asked: How would you count 5xx errors by client and path in one pass? · What are NR, NF and $NF? · Why does awk -F: '$3 >= 1000' on /etc/passwd also print nobody?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.