OnCallReady

Lesson 10.53 · Images & Builds · 9 min read

Image anatomy: manifest, config, layers, digests

In plain words

Imagine a flat-pack wardrobe delivered as a box. Inside there is a packing list, an instruction sheet, and several sealed bags of parts. The packing list names every bag and the instruction sheet by a unique fingerprint, so if anyone swaps a screw, the fingerprint on that bag no longer matches and you know. Two wardrobes that share the same bag of hinges only need one copy of that bag in the warehouse.

A container image is that box. The manifest is the packing list, the config is the instruction sheet (entrypoint, user, environment, history), and the layers are the bags, tar archives of what each build step changed. Everything is referenced by sha256 digest. docker save lets you open the box and look.

The problem

You have been using three different "IDs" for an image - the IMAGE ID from docker images, the sha256: digest from a pull or push, and the layer IDs in docker inspect - and they never match each other. You have also pulled images where only some layers downloaded. Both make sense once you see what an image is made of on the registry and on disk.

What you need to know already: tags vs digests (10.5), layers and whiteouts (10.10), docker save and tar (10.42), jq (Ch 7). A hash (sha256 here) is a fixed-length fingerprint computed from content: change one byte and the hash changes completely.

Three kinds of object

An image in a registry is not one file. It is a few files, each called a blob (just "a chunk of bytes, named by its hash"):

index      (optional) a list of manifests, one per platform - multi-arch images
manifest   the list of layer blobs and the config blob, each by digest
config     JSON: architecture, os, the runtime config (Env, Entrypoint, Cmd,
           User, WorkingDir...), the history, and rootfs.diff_ids
layers     tar archives (compressed in the registry), one per filesystem step

Everything references everything else by sha256 digest, which is why the whole structure is tamper-evident (any change is detectable): change one byte in one layer and its digest changes, so the manifest changes, so the manifest's digest - the thing you pinned - no longer matches.

Which ID is which:

Looking at it on disk

docker save -o file.tar <image> writes the image in the standard OCI layout (OCI = the Open Container Initiative, which standardises the format):

$ docker save -o orders.tar orders:tiny && mkdir o && tar -xf orders.tar -C o
$ ls o
blobs  index.json  manifest.json  oci-layout  repositories
$ cat o/manifest.json
[{"Config":"blobs/sha256/1f3e...","RepoTags":["orders:tiny"],"Layers":["blobs/sha256/da20...","blobs/sha256/5c1e...","blobs/sha256/77ab..."]}]

The config blob is plain JSON, so jq reads it. This command picks the config path out of manifest.json ($(...), Ch 6) and prints a few fields:

$ jq '{architecture, os, user: .config.User, entrypoint: .config.Entrypoint, history: [.history[].created_by]}' o/$(jq -r '.[0].Config' o/manifest.json)
{
  "architecture": "arm64",
  "os": "linux",
  "user": "app",
  "entrypoint": ["java","-jar","app.jar"],
  "history": [
    "# debian.sh --arch 'arm64' out/ 'bookworm' '@1726444800'",
    "CMD [\"bash\"]",
    "ENV JAVA_HOME=/opt/java/openjdk",
    ...
  ]
}

That config - history included - travels with the image to every registry and server. It is where docker history gets its data, and where a leaked ARG value lives.

Each layer blob is a tar of only what that step changed, including whiteouts for deletions. tar -tvf lists it (-t list, -v with owner, size and date, -f this file); .[0].Layers[-1] in jq is the last (newest) layer:

$ tar -tvf o/$(jq -r '.[0].Layers[-1]' o/manifest.json)
-rw-r--r-- 999/999    22012344 2026-09-23 10:00 app/app.jar

Why this matters day to day

What you can now do

Why it helps

This is the model behind several things you will do for real. Pinning image@sha256:... in a Dockerfile or a deploy config pins every layer, the entrypoint and the user, which is why it is the answer to "the same tag behaved differently yesterday". Scanners and image signatures also attach their results to the manifest digest. When a scanner reports a leaked secret, docker save and jq on the config blob show you exactly where it lives. When deploys are slow, knowing that pulls are per layer and registries store each layer once by digest tells you to keep bases stable and top layers small. And it explains the confusing mismatch between the IDs in docker pull output and docker inspect.

Commands in this lesson

docker ls cat jq tar

FAQ

What is the difference between the IMAGE ID and the image digest?

With Docker's classic image store, the IMAGE ID is the digest of the config JSON, computed locally, and the digest (name@sha256:..., printed by docker push and shown in RepoDigests) is the digest of the manifest, or of the index for multi-arch images. The digest is what a registry knows and what you pin. Note that with the containerd image store, the default on fresh Docker Engine 29 installs, docker images shows the manifest or index digest as the ID instead.

Why do the layer IDs in docker pull not match docker inspect?

The registry stores layers compressed, and docker pull shows the digests of those compressed blobs. docker inspect shows RootFS.Layers, which are diff IDs: the digests of the uncompressed tar archives. Same content, different bytes, different digests. The config's rootfs.diff_ids lists the uncompressed ones, the manifest lists the compressed ones.

What is an image index?

An index, also called a manifest list, is a small JSON document that points at one manifest per platform, such as linux/amd64 and linux/arm64. A multi-arch tag like nginx:1.27 points at an index, and the client picks the manifest for its own architecture. Pinning the index digest keeps the image multi-arch; pinning one platform's manifest digest ties it to that architecture.

Does pinning a digest also pin the entrypoint and environment?

Yes. The manifest references the config blob by digest, and the config contains the Entrypoint, Cmd, Env, User and WorkingDir, as well as the layer list. Change any of them and the config digest changes, so the manifest changes, so the digest you pinned no longer matches. A digest identifies the whole image, not just its files.

Where does docker history get its data?

From the history array in the config blob, which records the created_by instruction and whether each step produced a layer. It travels with the image to every registry and server. That is why an ARG value used in a RUN appears in docker history --no-trunc: the build recorded it there, and anyone who can pull the image can read it.

In an interview Junior

What is inside a container image, technically?

A few blobs, each named by its sha256 digest:

Everything references everything else by digest, so it is tamper-evident. Which ID is which: the image digest you pin is the manifest's digest; the IMAGE ID in docker images is the config's; the RootFS.Layers in docker inspect are digests of the uncompressed layer tars.

You can look: docker save -o x.tar IMAGE, tar -xf, then jq on manifest.json and the config blob. Pulls are per layer, so a server only downloads the layers it does not have.

Also asked: Why should you deploy by digest rather than by tag? · Why do the layer IDs in docker pull output not match docker inspect? · Why does pulling a new version of an image download only some layers?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.