Text Processing & jq: interview questions
The question you are most likely to get for each topic, a model answer, and what else comes up. From chapter 7 of the course.
Give me a one-liner for the top 10 client IPs in an nginx access log. Junior
awk '{print $1}' access.log | sort | uniq -c | sort -rn | head -10
awk '{print $1}'- the client IP is the first field; awk splits on runs of spaces.sort- puts identical IPs next to each other. Required:uniqonly merges adjacent duplicates.uniq -c- one line per IP, with its count in front.sort -rn- by that count, as a number, biggest first (plainsortwould put 9 above 10).head -10- the top ten.
The same four stages answer every "most common X" question by changing the field: $9 status codes, $7 paths. On a big log, awk can count in one pass: awk '{c[$1]++} END {for (ip in c) print c[ip], ip}' access.log | sort -rn | head. And check the field number against one real line first (head -1 access.log).
Also asked: From a JSON array, how do you print two fields as TSV, filtered by a third? · How do you replace a string in a config file in place, safely? · How do you see the lines just before and after an error in a log?
How do you find an error in a big log and see what led up to it? Junior
grep prints the lines that match a pattern; the flags do the rest:
grep -B2 -A2 'upstream timed out' error.log- 2 lines Before and After each match (-C2both). The match says what, the context usually says why - the most useful grep in an incident.grep -c ' 502 ' access.log- just the count of matching lines (the spaces stop it matching a byte count like 5021).-nline numbers,-iignore case,-vlines that do not match,-wwhole words.grep -oE '^[0-9]+\.[0-9]+\.[0-9]+\.[0-9]+' access.log--oprints only the matching part, here the IP, ready forsort | uniq -c.-Ffor a fixed string, no regex;-qforiftests (no output, only the exit status).
Always single-quote the pattern, so the shell does not expand * or $ first. The error log is root-only: sudo grep.
Also asked: How do you search a whole directory and list only the files that contain a string? · How would you use grep inside an if in a script? · What is the difference between grep -c and grep ... | wc -l?
Learn it: 7.1 grep
What is the difference between basic and extended regular expressions? Junior
The operators are the same; which characters need a backslash is reversed.
- BRE (basic) - the default in
grepandsed:+ ? | ( ) { }are literal characters, and you backslash them to make them operators:a\{2,\}. - ERE (extended) -
grep -E,sed -E, and awk always: they are operators:a{2,},(dev|prod); backslash them for the literal character. - PCRE (
grep -P) adds\d,\w, lazy*?and lookarounds.
., *, ^, $ and [...] mean the same in all of them. I use -E everywhere so a pattern reads the same in grep, sed and awk. Two classic bugs: \d in grep -E is a literal d, not a digit (write [0-9]); and an unescaped . matches any character, so 10.0.4.17 also matches 110.0.4.170 - escape the dots and anchor it, or use grep -F.
Also asked: Write a regex for an IPv4 address in a log line. What are the pitfalls? · What does "greedy" mean, and how do you match one quoted field? · What are capture groups, and how do you use them in sed?
How do you find the most common HTTP status codes in an access log? Junior
awk '{print $9}' access.log | sort | uniq -c | sort -rn
In nginx's log the status code is the ninth space-separated field, so awk '{print $9}' extracts it. sort groups identical codes together - needed because uniq has no memory and only collapses adjacent duplicates. uniq -c counts each run. sort -rn orders by that count, numerically (-n) and biggest first (-r); without -n, "9" would sort above "10".
Check the field number against one real line first - a different log format shifts every field. And cut -d' ' -f9 is not a replacement for awk here: cut treats every single space as a delimiter, so runs of spaces give empty fields.
Also asked: What is the difference between sort | uniq and sort -u? · Why does uniq almost always come after sort? · When does cut fail where awk works?
Learn it: 7.4 cut, sort, uniq, tr, column
How do you replace a string in a file in place, safely? Junior
sed -i.bak 's/old/new/g' file: -i writes the result back into the file, .bak keeps the original as file.bak, and g replaces every match on a line, not just the first.
"Safely" means three things:
- Run it without
-ifirst and read the output - sed has no undo. - Keep a backup:
-i.bakon anything under/etc, e.g.sudo sed -i.bak 's/read.timeout.ms=0/read.timeout.ms=5000/' /etc/orders/app.conf. - Make the pattern specific: anchor it (
^), escape dots, so it cannot hit a comment that mentions the same key.
If the pattern contains slashes, use another delimiter: s|/old/path|/new/path|. And sed -i needs write permission on the directory, because it writes a temporary file and renames it.
Also asked: How do you print only lines 40 to 60 of a file? · How do you strip comments and blank lines from a config file? · When would you use something other than sed?
Learn it: 7.6 sed
How do you print one column of a command's output, and how would you add a column up? Junior
awk '{print $2}' - awk reads line by line, splits each line into fields on runs of whitespace, and $2 is the second field ($1 the first, $NF the last). A pattern in front makes it a filter: awk '$9 ~ /^5/ {print $1}' access.log prints the client of every 5xx. For another delimiter, -F: awk -F: '{print $1}' /etc/passwd.
Adding up uses an accumulator and an END block, which runs once after the last line:
awk '{sum += $10} END {print sum}' access.log total bytes
awk '{s += $NF} END {printf "%.3f\n", s/NR}' access.log average time
awk '{c[$1]++} END {for (ip in c) print c[ip], ip}' access.log count per IP
Variables start at 0 without being declared; NR in END is the number of lines read. Keep the program in single quotes so the shell leaves $1 alone.
Also asked: How would you count 5xx errors by client and path in one pass? · What are NR, NF and $NF? · Why does awk -F: '$3 >= 1000' on /etc/passwd also print nobody?
Learn it: 7.8 awk
From a JSON array of objects, how do you print two fields as TSV, filtered by a third? Junior
With pods.json, namespace and name of every entry that is not Running:
jq -r '.items[] | select(.status.phase != "Running") | [.metadata.namespace, .metadata.name] | @tsv' pods.json
.items[]- each element of the array as its own value (.itemsalone would be one array).select(...)- keep only the elements where the condition is true; it is grep for structured data.[a, b] | @tsv- build an array of the two fields and format it as tab-separated values (@csvfor commas).-r- raw output: plain text without JSON quotes, so the next command in a pipe can use it.
To explore the shape first: jq '.items[0]'. If the filter value comes from the shell, pass it in with --arg instead of pasting it into the program: jq --arg ns payments '.items[] | select(.metadata.namespace == $ns)'.
Also asked: What is the difference between .items and .items[]? · Why do you almost always want jq -r? · How do you find a key at any depth in a JSON document?
Learn it: 7.11 jq
How do you count the entries per group (per namespace) in a JSON document with jq? Junior
jq -c '[.items[] | {ns: .metadata.namespace}] | group_by(.ns) | map({ns: .[0].ns, n: length})' pods.json
Read it in steps: [ ... ] collects one small object per entry into a single array; group_by(.ns) sorts that array and splits it into arrays of equal .ns; map({...}) turns each group into one object - .[0].ns is the group's key and length its size. Result: [{"ns":"default","n":2},...,{"ns":"payments","n":4}]. It is sort | uniq -c for JSON.
Related tools from the same family: add sums an array (add // 0 for an empty one), unique gives distinct values, max_by(f) the element with the largest f. And to use a result in a script: jq -e sets the exit status from the last output (1 for null/false, 4 for no output), so it works in an if.
Also asked: How do you list or filter the keys of an object you do not know in advance, like labels? · Why does --arg n 2 with > $n never match, and what do you use instead? · How do you change one value in a JSON file without destroying it?
Learn it: 7.13 jq beyond select
What does xargs do, and how do you use it safely with filenames? Junior
Many commands (rm, chmod, systemctl) take their list as arguments, not from stdin. xargs CMD reads words from stdin and runs CMD with them added as arguments: echo a b c | xargs echo runs echo a b c.
The danger: xargs splits on spaces, so a file called my report.log arrives as two arguments. The safe idiom separates names with a NUL byte, the one character that cannot be in a filename:
find . -name '*.log' -print0 | xargs -0 rm
Useful flags: -r do nothing if the input is empty, -n1 one argument per run, -I{} put each line where {} is, -P4 four at a time in parallel (effective, and good at overwhelming whatever is on the other end). For a list that comes from find, -exec rm {} + or -delete does the same without a pipe.
Also asked: What is the difference between find -exec cmd {} \; and -exec cmd {} +? · Why add -r to xargs? · When would you use xargs -I{}?
Learn it: 7.16 xargs
Practise these answers with flashcards and labs Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.