Why this lesson
"Total bytes served", "average response time", "5xx count per client and path" - questions that need columns and arithmetic. grep only knows whole lines; cut cannot add. awk is a small programming language made for exactly this: it reads a file line by line, splits each line into fields, and lets you compare, add up and count them.
What you need to know already: fields and delimiters (6.8); regex (7.3); associative arrays (6.12) - awk has them too; printf formats (6.5); the access log's fields (7.1, 7.4).
awk is "for each line, if PATTERN then ACTION"
On your VM: Ubuntu's
awkis mawk, not gawk (two versions of the same tool). Everything in this chapter works in both; gawk-only extras (PROCINFO["sorted_in"],strftime,gensub) needsudo apt install gawk.
awk 'PATTERN { ACTION }' file
The text in quotes is an awk program. For every line (awk calls it a record), awk checks the PATTERN; if it is true, it runs the ACTION in braces.
Leave out the pattern and the action runs on every line. Leave out the action and awk prints the matching line. So awk '/error/' works like grep, and awk '{print $2}' like a smarter cut.
The variables
awk fills these in for you on every line:
$0 the whole line NF number of fields on this line
$1.. the fields NR record (line) number, running total
$NF the LAST field FS input field separator (default: whitespace)
$(NF-1) second to last OFS output field separator
$1 here is awk's own "field 1", not a shell variable - that is why the awk program is always in single quotes, so the shell leaves $1 alone.
awk splits on runs of whitespace by default, which is precisely what cut cannot do. For any other delimiter, -F (field separator):
awk -F: '{print $1}' /etc/passwd
awk -F, '{print $3}' data.csv
awk 'BEGIN{FS=":"; OFS="\t"} {print $3, $1}' /etc/passwd
(BEGIN{...} runs once before the first line - below. OFS is what print a, b puts between a and b; \t is a tab.)
Patterns
~ means "matches this regex", !~ "does not match", == "is equal to":
awk '$3 >= 1000' numeric comparison on a field
awk '$9 ~ /^5/' field matches a regex
awk '$9 !~ /^2/' does not match
awk '/GET/ && $9 == 200' combine
awk 'NR > 1' skip the header
awk 'NR==10, NR==20' a range of lines
BEGIN, END and accumulation
BEGIN { ... } runs once before any line is read; END { ... } runs once after the last line. An accumulator is a variable you add to on every line and print in END. awk variables start at 0 (or empty) without being declared, and sum += $10 means "add field 10 to sum". This is where awk stops being "cut with extras":
awk '{sum += $10} END {print sum}' access.log total bytes
awk '{s += $NF} END {printf "%.3f\n", s/NR}' access.log average, 3 decimals
awk '{c[$1]++} END {for (ip in c) print c[ip], ip}' access.log
The second divides by NR, which in END is the number of lines read. %.3f is a printf format (6.5): a number with 3 decimals.
The last one is the important one. c[$1]++ uses an associative array (6.12
- in awk they need no declaring) keyed by the first field, adding 1 each time it
sees that value. for (ip in c) walks every key in END. That is sort | uniq -c in a single pass, and it can count anything - including combinations:
awk '$9 ~ /^5/ {c[$1" "$7]++} END {for (k in c) print c[k], k}' access.log \
| sort -rn | head
"For every 5xx, count by client and path." $1" "$7 glues field 1, a space and field 7 into one string - a composite key. (The trailing \ continues the command on the next line.) Two greps and a cut could not do that.
printf in awk works like the shell's: printf "%-20s %6.2f\n", name, value - but with commas between the arguments.
awk versus the pipeline
sort | uniq -c | sort -rn is easier to type and easier to read. awk's array version is one pass instead of three, which matters on a large log, and lets you key on something computed rather than a whole field.
Use whichever is clearer until the log is big enough for it to matter.
What you can now do
- Filter lines on a field (
$9 ~ /^5/,$3 >= 1000) and print chosen fields. - Sum, average and count with accumulators and an END block.
- Count by one or two fields at once with
c[key]++.