OnCallReady

Lesson 10.10 · Images & Builds · 18 min read

Reading docker history: where the bytes went

In plain words

Imagine a cake made of layers, where each layer is baked separately and glued on top. If you decide the third layer has too much chocolate, you cannot scrape it out; you can only put a new layer on top with a note that says 'pretend the chocolate is not there'. The cake still weighs the same, plus the note.

docker history lists the cake's layers from the top down, with the bytes each instruction added. Deleting in a later RUN only adds a whiteout, so the bytes stay. Changing a file's owner copies the whole file into a new layer. So you read history to find the fat row, then fix the Dockerfile so the mess is cleaned up in the same step that made it.

Why this matters

"Why is our image 1GB?" Every server that runs it downloads that gigabyte, and every extra package is more to keep patched. docker history tells you which Dockerfile line added which bytes, so you fix the right line instead of guessing.

What you need to know already: layers and overlayfs ("What a container actually is"), the Dockerfile instructions (previous lesson), chown and file owners (Ch 4).

history is a list of instructions and their sizes

$ docker history orders:naive
IMAGE          CREATED          CREATED BY                                      SIZE      COMMENT
cf3ab9a158c8   2 minutes ago    ENTRYPOINT ["/bin/sh" "-c" "java -jar target…   0B        buildkit.dockerfile.v0
<missing>      2 minutes ago    RUN /bin/sh -c mvn package -DskipTests # bui…   275MB     buildkit.dockerfile.v0
<missing>      2 minutes ago    COPY . . # buildkit                             112MB     buildkit.dockerfile.v0
<missing>      2 minutes ago    WORKDIR /app                                    0B        buildkit.dockerfile.v0
<missing>      3 weeks ago      CMD ["mvn"]                                     0B        buildkit.dockerfile.v0
<missing>      3 weeks ago      ENTRYPOINT ["/usr/local/bin/mvn-entrypoint.s…   0B        buildkit.dockerfile.v0
<missing>      3 weeks ago      COPY /usr/share/maven /usr/share/maven # bui…   10.4MB    buildkit.dockerfile.v0
<missing>      3 weeks ago      RUN |1 MAVEN_VERSION=3.9.9 /bin/sh -c apt-ge…   58.1MB    buildkit.dockerfile.v0
<missing>      3 weeks ago      RUN /bin/sh -c set -eux;     ARCH="$(dpkg --…   350MB     buildkit.dockerfile.v0
...
<missing>      3 weeks ago      ADD file:3a1c4e2b7d5f in /                      101MB

How to read it:

Here the story is complete in four rows: 101MB of Ubuntu, 350MB of JDK, 112MB of COPY . . (the whole project directory, git history included) and 275MB from the Maven build (the downloaded dependencies plus target/). The application JAR is 22MB of that.

docker history --no-trunc orders:naive          full commands
docker history --format '{{.Size}}\t{{.CreatedBy}}' orders:naive
docker history -H=false orders:naive            sizes in bytes, for sorting

--format picks columns with a Go template (\t is a tab); -H=false turns off human-readable sizes so sort -n (Ch 7) works on them.

The first rule of layers: deleting later frees nothing

Layers are stacked, not merged. A later layer can only record that a file is gone, with a whiteout entry - a marker file that hides it. The bytes stay in the earlier layer, and every pull downloads them:

RUN apt-get update && apt-get install -y build-essential    # +270MB
RUN apt-get purge -y build-essential && apt-get autoremove -y
RUN rm -rf /var/lib/apt/lists/*                             # +0B

(build-essential is the C compiler package; purge removes a package and its config; /var/lib/apt/lists is the package index apt-get update downloads.)

<missing>   RUN /bin/sh -c rm -rf /var/lib/apt/lists/* # b…   0B
<missing>   RUN /bin/sh -c apt-get purge -y build-essent…     1.2kB
<missing>   RUN /bin/sh -c apt-get update && apt-get ins…     317MB

The image is 317MB bigger, forever. The fix is to create and delete in the same RUN, so the temporary files never reach a layer:

RUN apt-get update \
 && apt-get install -y --no-install-recommends curl ca-certificates \
 && rm -rf /var/lib/apt/lists/*

--no-install-recommends matters too: Ubuntu installs "recommended" extra packages by default, often doubling what you install. The Debian and Ubuntu base images already delete downloaded .deb files automatically (/etc/apt/apt.conf.d/docker-clean), so apt-get clean adds nothing; the package lists are what you must remove.

The second rule: changing metadata copies the file

overlayfs cannot change a file's owner or mode in a lower layer. It copies the whole file up into the new layer (a copy-up) and changes it there:

COPY target/app.jar /app/app.jar       # +22MB
RUN chown -R 1000:1000 /app            # +22MB again
<missing>   RUN /bin/sh -c chown -R 1000:1000 /app # b…   22MB
<missing>   COPY target/app.jar /app/app.jar # buildkit    22MB

Your JAR is now stored twice. With a 400MB folder of dependencies it is 400MB twice. Set ownership as you copy:

COPY --chown=1000:1000 target/app.jar /app/app.jar

(and --chmod=755 exists for modes, Ch 4). Same result, one copy.

The third rule: COPY'd files belong to root

Whatever USER says, COPY and ADD create files owned by root unless you pass --chown. So this common Dockerfile

USER 1000
COPY . /app

produces an app that cannot write to its own directory:

# app:1 = an image built from that Dockerfile
docker run --rm app:1 touch /app/cache.db
touch: cannot touch '/app/cache.db': Permission denied

USER changes who runs things; it does not change who owns things copied in.

What to look for, in order

When an image is too big, run docker history --no-trunc and look for:

  1. A huge base at the bottom - a JDK, node:22, python:3.12 (the full variants are around 1GB). Switch to a smaller runtime base (later lessons).
  2. A huge COPY - usually COPY . . with no .dockerignore: .git, target/, node_modules/ (Node's dependency folder).
  3. A RUN that downloads (package managers, dependency caches) whose cleanup happens in a different RUN.
  4. The same size twice - a chown -R or chmod -R after a COPY.
  5. Build tools in the final image - compilers, Maven, -dev packages.

The history tells you which row; the Dockerfile tells you why.

What you can now do

Why it helps

The platform team asks why the payments image is 1.1GB and servers take a minute to pull it whenever one is added. docker history --no-trunc answers in one screen: a 350MB JDK base, a 112MB COPY . . that dragged .git in, a 275MB Maven cache. It is also the first command in a security incident: if someone reports a token in an image, the ARG values and full RUN lines are right there in the history, visible to anyone who can pull. In reviews you will spot the classic mistakes by eye once you have seen their rows: a cleanup-only RUN rm -rf, which frees nothing, and a chown -R after COPY, which doubles the size. You stop guessing and point at the row.

Commands in this lesson

docker

FAQ

Why do most rows show <missing> in the IMAGE column?

It is normal. BuildKit does not keep an intermediate image for each step, only the final image, so only the top row has an ID. Older builds with the legacy builder showed intermediate IDs. <missing> does not mean a layer is missing; the data is all there.

I deleted the files in a later RUN. Why is the image still big?

Layers stack; they are never merged. A later layer can only record a whiteout saying the file is gone, and the bytes stay in the earlier layer, which every pull still downloads. Create and delete in the same RUN, for example apt-get update && apt-get install ... && rm -rf /var/lib/apt/lists/*, so the temporary files never reach a snapshot.

Why does chown -R after COPY double the size?

overlayfs cannot change the owner or mode of a file in a lower layer in place. It copies the whole file up into the new layer and changes it there. So a 22MB jar becomes 44MB, and a 400MB node_modules becomes 800MB. Use COPY --chown=1000:1000 and --chmod to set ownership and modes as the files are copied.

I set USER 1000 before COPY. Why are the files still owned by root?

USER changes who runs RUN, CMD and ENTRYPOINT. It does not change who owns files that COPY or ADD create; those are root-owned unless you pass --chown. So an app running as 1000 cannot write into a directory you copied in. Either COPY --chown, or create the writable directory and chown it in a RUN before switching USER.

Does apt-get clean make my image smaller?

On the official Debian and Ubuntu base images, no. They already ship an apt config (docker-clean) that deletes downloaded .deb files automatically. What you must remove is the package index in /var/lib/apt/lists/, in the same RUN as the install. --no-install-recommends usually saves more than anything else, because recommended packages can double the install.

In an interview Junior

How do you find out why a Docker image is so large?

docker history --no-trunc IMAGE lists every instruction with the bytes it added, newest at the top; the base image is the bottom rows. Look for, in order:

  1. A huge base - a JDK, node:22, a full python image.
  2. A huge COPY - usually COPY . . without a .dockerignore (.git, target/, node_modules/).
  3. A RUN that downloads whose cleanup is in a different RUN. Deleting in a later layer frees nothing: it only adds a whiteout; the bytes stay in the earlier layer. Create and delete in the same RUN (apt-get update && apt-get install -y --no-install-recommends ... && rm -rf /var/lib/apt/lists/*).
  4. The same size twice - chown -R after a COPY makes a copy-up of every file; use COPY --chown.
  5. Build tools in the final image - compilers, Maven.

The history says which row; the Dockerfile says why.

Also asked: Why does RUN rm -rf on its own line not make the image smaller? · Why are files you COPY owned by root even after USER? · How would you check whether a password was written into an image's history?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.