Why this lesson
More and more tools print JSON instead of plain lines - and most of what you will query later in this course speaks it. JSON is nested, so grep, cut and awk fall apart on it: the value you want is three levels deep, and the same key appears in many places. jq is the tool that understands the structure.
What you need to know already: pipes (1.7); sort -u and uniq (7.4); the difference between single and double quotes (6.6).
JSON in five minutes
JSON (JavaScript Object Notation) is a text format for structured data. It has just a few building blocks:
"text" a string
42 3.5 a number
true false null true/false, and "nothing"
[1, 2, 3] an ARRAY: an ordered list of values
{"name": "web", "n": 2} an OBJECT: a set of key: value pairs
A key is the name on the left of a :, the value is what follows. Values nest: an object can hold arrays of objects, and so on.
The file you will use
~/labs/data/pods.json is a JSON list of running programs, exported from a group of machines that run as one. Here it is just JSON data: one entry per running program. The shape, trimmed:
{
"items": [ <- an array, one object per pod
{
"metadata": {"name": "web-7d9f6c5b8-2xkpq", "namespace": "default"},
"spec": {"containers": [{"image": "nginx:1.27"}]},
"status": {"phase": "Running"}
},
...
]
}
name is the entry's name; namespace is the group it is filed under (like a folder: payments, monitoring); phase is its state (Running, Pending = waiting to start, Failed, Succeeded = finished). image names the packaged program it runs, with a version after the :.
Later (Ch 15): these entries are Kubernetes pods, and this is the JSON
kubectl get pods -o jsonprints. Everything you learn here works on it as is.
Everything is a filter
jq reads JSON, applies a filter (a small program describing what to take out), and prints JSON. | inside the filter chains filters, just like a shell pipe. . is "the whole input"; .key is "the value of this key"; .key[] is "each element of this array, one at a time".
jq '.' file.json pretty-print (and check it is valid JSON)
jq '.items' file.json one key's value
jq '.items[]' file.json each ELEMENT of the array, separately
jq '.items[].metadata.name' that key from each element
jq '.items | length' how many elements
The difference between .items and .items[] is the thing to get straight: the first gives you one array, the second gives you each element as its own output. Everything after it operates per element.
-r, and why you always want it
$ cd ~/labs/data
$ jq '.items[0].metadata.name' pods.json
"web-7d9f6c5b8-2xkpq"
$ jq -r '.items[0].metadata.name' pods.json
web-7d9f6c5b8-2xkpq
Without -r (raw output) strings come out as JSON strings, with quotes - which is correct, and useless the moment you pipe it into anything else.
select: the filter
select(CONDITION) keeps the values for which CONDITION is true. == is equal, != not equal:
jq '.items[] | select(.status.phase != "Running")'
jq '.items[] | select(.status.containerStatuses[0].restartCount > 3) | .metadata.name'
jq '.items[] | select(.metadata.namespace == "payments")'
(restartCount is how many times the program was restarted after dying; it sits inside status.containerStatuses, an array, and [0] takes its first element.)
select passes through the values for which its condition is true and drops the rest. It is grep for structured data, and it is most of what you will use jq for.
Building output
jq -r '.items[] | [.metadata.namespace, .metadata.name] | @tsv'
jq -r '.items[] | "\(.metadata.name) is \(.status.phase)"'
jq -c '.items[] | {name: .metadata.name, phase: .status.phase}'
[a, b] | @tsv is the one to remember: build an array, format it as TSV (tab-separated values, 6.27). @csv is the comma sibling. "\(expr)" is string interpolation: put the value of expr inside a string. {name: ...} builds a new object, and -c prints each result compactly on one line.
Aggregating
To aggregate - sum, count distinct values - first collect everything into one array with [ ... ], then apply add (sum), unique (distinct values), group_by (split into groups) or sort_by. map(f) applies f to every element of an array.
jq '[.items[].status.containerStatuses[]?.restartCount] | add' sum
jq '.items | map(.status.phase) | unique' distinct values
jq '.items | group_by(.status.phase) | map({phase: .[0].status.phase, n: length})'
jq '.items | sort_by(.status.containerStatuses[0].restartCount) | reverse | .[0:3]'
.[0:3] is a slice: elements 0, 1 and 2. reverse flips the order, so together they give the three with the most restarts.
Surviving missing keys
Not every object has every key (an entry that never started has no containerStatuses). Asking for a missing key gives null; asking for a key on something that is not an object is an error. Three tools:
jq '.items[] | .spec.nodeName // "not placed yet"' default if null/false
jq '.items[] | .metadata.labels?' ? = no error, just skip
jq '.. | .image? // empty' recursive descent
(nodeName is the machine the program was placed on; a Pending entry has none yet. []? in the sum above is the same ?: iterate if it is an array, skip if not.)
A // B means "A, or B if A is null or false". empty produces no output at all.
.. (recursive descent) walks every value at every depth. .. | .image? // empty means "find an image key anywhere in this document, and ignore everything that has not got one" - exactly right when the same key can sit at different depths (here, an entry can list its image in two different places).
Passing shell values in
jq --arg ns payments '.items[] | select(.metadata.namespace == $ns)'
--arg NAME VALUE makes $NAME available inside the filter. (That $ns is jq's, not the shell's - the single quotes keep the shell away from it.) Never paste a shell variable into the jq program text; --arg passes the value in safely and quotes it for you.
What you can now do
- Read a JSON document's shape and pull out a nested value.
- Filter with
select, and print plain text with-ror TSV with@tsv. - Find a key at any depth with
.. | .key? // empty.