OnCallReady

Lesson 7.13 · Text Processing & jq · 19 min read

jq beyond select

In plain words

Last lesson you learned to open boxes and pick things out. Now you learn to rearrange the boxes. Imagine a bag of Lego: group_by sorts the bricks into piles by colour, map does the same thing to every pile, length counts a pile. to_entries takes a box with labelled drawers and lays the drawers out in a row so you can look at each one; from_entries puts them back. |= swaps one brick in place and hands you back the whole model.

In jq terms: exploring an unknown document with keys and length, counting with group_by and add, iterating labels with to_entries, editing with |=, and reading many JSON documents at once with -s.

Why this lesson

select and @tsv answer "which ones". Real questions go further: "how many per group?", "what are its labels?", "change this one number and save a copy", "fail my script if any entry is on a latest version". This lesson is the rest of jq you will actually use.

What you need to know already: JSON, .key, .[], select, -r and --arg (7.11, 7.12); sort f > f truncating the file (6.17); if on an exit status (6.5).

Looking at an unknown document

keys lists an object's keys (sorted); join(",") glues an array of strings with commas; has("k") is true if the key exists:

$ cd ~/labs/data
$ jq -r '.items[0] | keys | join(",")' pods.json
apiVersion,kind,metadata,spec,status
$ jq '.items | length' pods.json
12
$ jq '.items[0].metadata | has("labels")' pods.json
true

keys (sorted) and length are how you explore: what fields are there, how many elements. jq -c '.items[0]' prints one element compactly on a line.

Objects as lists: to_entries

$ jq -c '.items[0].metadata.labels | to_entries' pods.json
[{"key":"app","value":"web"}]
$ jq -c '.items[0].metadata.labels | to_entries | from_entries' pods.json
{"app":"web"}
$ jq -r '.items[0].metadata.labels | to_entries[] | "\(.key)=\(.value)"' pods.json
app=web

to_entries turns an object into [{key, value}] so you can iterate, filter or format it; from_entries turns it back; with_entries(f) does both around f (with_entries(select(.key | startswith("app"))) keeps some keys).

Labels are free-form key: value tags people attach to things (app: web, team: payments). You do not know their keys in advance - which is exactly when you need to_entries.

Counting and grouping

$ jq -c '[.items[] | {ns: .metadata.namespace}] | group_by(.ns) | map({ns: .[0].ns, n: length})' pods.json
[{"ns":"default","n":2},{"ns":"kube-system","n":3},{"ns":"monitoring","n":3},{"ns":"payments","n":4}]

Read it right to left from the end: group_by(.ns) sorts and splits the array into arrays of equal .ns; map({...}) turns each group into one object; .[0].ns is the group's key and length its size. It is sort | uniq -c for JSON.

jq '[.items[].status.containerStatuses[]?.restartCount] | add'     sum: 14
jq '.items | map(.status.phase) | unique'                         distinct values
jq '.items | max_by(.status.containerStatuses[0].restartCount) | .metadata.name'
jq 'reduce .items[] as $p (0; . + ($p.spec.containers | length))'  a fold

max_by(f) picks the element with the largest f. reduce ITEMS as $x (START; UPDATE) is a fold: start with a value, and update it once per item - here, add up how many packaged programs (the containers list, 7.11) each entry runs.

add sums numbers, concatenates strings and arrays, and merges objects. It returns null for an empty array, so add // 0 when that matters.

Passing values in

jq --arg ns payments '.items[] | select(.metadata.namespace == $ns) | .metadata.name'
jq --argjson n 2 '.items[] | select(.status.containerStatuses[0].restartCount > $n)'

--arg always passes a string. --arg n 2 then > $n compares a number with the string "2" - and in jq every number is less than every string, so it is always false. Numbers and booleans go in with --argjson.

Changing a document

$ jq '.items[0].metadata.labels.app |= ascii_upcase | .items[0].metadata.labels' pods.json
{
  "app": "WEB"
}

A path here is the chain of keys and indexes to a value (.items[0].metadata.labels.app). |= updates the value at a path and returns the whole document. = assigns, += adds. ascii_upcase upper-cases a string. jq never edits a file: write to a new file and move it into place.

jq '.spec.replicas = 3' app.json > app.json.new && mv app.json.new app.json

jq ... > app.json on the same file truncates it before jq reads it - the same trap as sort f > f (6.17).

Several documents: -s and -n

$ echo '{"a":1}{"a":2}' | jq -s 'map(.a) | add'
3

-s (slurp) reads every input document into one array - how you aggregate over newline-delimited JSON: one JSON object per line, a common format for application logs. -n starts with null and no input; with --arg it builds JSON safely from shell values: jq -n --arg u "$USER" '{user: $u}'.

Matching strings

jq '.items[] | select(.metadata.name | test("^web")) | .metadata.name'
jq '.items[] | select(.metadata.name | startswith("refund"))'
jq -r '.items[].spec.containers[].image | split(":") | .[1]'      the tag

test() takes a PCRE-style regex (7.3) (test("^web-\\d"), note the doubled backslash inside a jq string inside shell quotes). split(":") cuts a string into an array at each : - so .[1] is the version after the colon. join, ascii_downcase, ltrimstr / rtrimstr (remove a prefix / suffix) cover most other string work.

Exit status and errors

$ jq -e '.items[] | select(.metadata.name=="nope")' pods.json; echo $?
4
$ jq '.items.metadata' pods.json
jq: error (at pods.json:545): Cannot index array with string "metadata"
$ jq '.items[' pods.json
jq: error: syntax error, unexpected end of file at <top-level>, line 1:
.items[
jq: 1 compile error

What you can now do

Why it helps

Real JSON is rarely one flat list. Labels have arbitrary keys, so "find all entries with a label starting with app" needs to_entries or with_entries. "How many per group" or "total restarts across everything" is group_by and add, the JSON version of sort | uniq -c. Application logs in JSON lines need -s to aggregate. Release scripts that bump a version in a JSON file need |= and the write-to-temp-then-mv pattern, and the bug where --arg turns a number into a string will one day make a threshold check silently never fire. When a script fails with "Cannot index array with string", this lesson is how you read that error in five seconds instead of fifteen minutes.

Commands in this lesson

cd jq echo

FAQ

What does "Cannot index array with string" mean?

You asked an array for a named field. .items.metadata tries to take metadata from the array itself, but arrays only have numeric indexes. You meant .items[].metadata, or .items[0].metadata for one element. The reverse error, "Cannot index string with string", means you went one level too deep, for example asking a name string for a field.

Why is my --arg comparison always false?

--arg always passes a string. --arg n 2 then select(.restartCount > $n) compares a number with the string "2", and jq's ordering puts every number below every string, so it is never true. No error, just no output. Use --argjson n 2 for numbers, booleans and JSON values, or convert inside the filter with ($n | tonumber).

What is the difference between map and .[]?

map(f) is shorthand for [.[] | f]: it iterates and collects the results back into an array. .[] | f iterates without collecting, producing separate outputs. Use map when the next step needs the array (map(.name) | unique, | length, | add); use .[] when each result goes out on its own line, for example into @tsv or xargs.

When do I need -s?

When the input is several JSON documents rather than one: newline-delimited JSON logs, several JSON files concatenated with cat, one object per line from an API. Normally jq runs the filter once per input document. -s (slurp) reads them all into a single array first, so you can map, group_by or add across them. On huge inputs that uses memory; for simple per-line filtering, leave -s off.

Can jq edit a file in place?

No, jq has no -i. It always writes to stdout. The pattern is jq '.spec.replicas = 3' f.json > f.json.new && mv f.json.new f.json. Redirecting straight back to f.json truncates the file before jq reads it and leaves you with an empty file, the same trap as sort f > f. sponge from moreutils is an alternative if installed.

In an interview Junior

How do you count the entries per group (per namespace) in a JSON document with jq?

jq -c '[.items[] | {ns: .metadata.namespace}] | group_by(.ns) | map({ns: .[0].ns, n: length})' pods.json

Read it in steps: [ ... ] collects one small object per entry into a single array; group_by(.ns) sorts that array and splits it into arrays of equal .ns; map({...}) turns each group into one object - .[0].ns is the group's key and length its size. Result: [{"ns":"default","n":2},...,{"ns":"payments","n":4}]. It is sort | uniq -c for JSON.

Related tools from the same family: add sums an array (add // 0 for an empty one), unique gives distinct values, max_by(f) the element with the largest f. And to use a result in a script: jq -e sets the exit status from the last output (1 for null/false, 4 for no output), so it works in an if.

Also asked: How do you list or filter the keys of an object you do not know in advance, like labels? · Why does --arg n 2 with > $n never match, and what do you use instead? · How do you change one value in a JSON file without destroying it?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.