Why this matters
"Why is our image 1GB?" Every server that runs it downloads that gigabyte, and every extra package is more to keep patched. docker history tells you which Dockerfile line added which bytes, so you fix the right line instead of guessing.
What you need to know already: layers and overlayfs ("What a container actually is"), the Dockerfile instructions (previous lesson), chown and file owners (Ch 4).
history is a list of instructions and their sizes
$ docker history orders:naive
IMAGE CREATED CREATED BY SIZE COMMENT
cf3ab9a158c8 2 minutes ago ENTRYPOINT ["/bin/sh" "-c" "java -jar target… 0B buildkit.dockerfile.v0
<missing> 2 minutes ago RUN /bin/sh -c mvn package -DskipTests # bui… 275MB buildkit.dockerfile.v0
<missing> 2 minutes ago COPY . . # buildkit 112MB buildkit.dockerfile.v0
<missing> 2 minutes ago WORKDIR /app 0B buildkit.dockerfile.v0
<missing> 3 weeks ago CMD ["mvn"] 0B buildkit.dockerfile.v0
<missing> 3 weeks ago ENTRYPOINT ["/usr/local/bin/mvn-entrypoint.s… 0B buildkit.dockerfile.v0
<missing> 3 weeks ago COPY /usr/share/maven /usr/share/maven # bui… 10.4MB buildkit.dockerfile.v0
<missing> 3 weeks ago RUN |1 MAVEN_VERSION=3.9.9 /bin/sh -c apt-ge… 58.1MB buildkit.dockerfile.v0
<missing> 3 weeks ago RUN /bin/sh -c set -eux; ARCH="$(dpkg --… 350MB buildkit.dockerfile.v0
...
<missing> 3 weeks ago ADD file:3a1c4e2b7d5f in / 101MB
How to read it:
- Newest at the top, the base image's first layer at the bottom. Your Dockerfile is the top few rows; everything from "3 weeks ago" down is the base image (someone else's Dockerfile).
- IMAGE is only filled for the top row.
<missing>is normal - BuildKit does not keep an image per step, only the final one. - CREATED BY - the instruction. It is cut off;
--no-truncprints the whole command - and that is the first command to run when you suspect a password leaked into an image, because build args used by a RUN are recorded right there:RUN |1 MAVEN_VERSION=3.9.9 /bin/sh -c ...means one ARG was set. - SIZE - the bytes that instruction added.
0Brows are metadata. - COMMENT - which builder made the row; ignore it.
Here the story is complete in four rows: 101MB of Ubuntu, 350MB of JDK, 112MB of COPY . . (the whole project directory, git history included) and 275MB from the Maven build (the downloaded dependencies plus target/). The application JAR is 22MB of that.
docker history --no-trunc orders:naive full commands
docker history --format '{{.Size}}\t{{.CreatedBy}}' orders:naive
docker history -H=false orders:naive sizes in bytes, for sorting
--format picks columns with a Go template (\t is a tab); -H=false turns off human-readable sizes so sort -n (Ch 7) works on them.
The first rule of layers: deleting later frees nothing
Layers are stacked, not merged. A later layer can only record that a file is gone, with a whiteout entry - a marker file that hides it. The bytes stay in the earlier layer, and every pull downloads them:
RUN apt-get update && apt-get install -y build-essential # +270MB
RUN apt-get purge -y build-essential && apt-get autoremove -y
RUN rm -rf /var/lib/apt/lists/* # +0B
(build-essential is the C compiler package; purge removes a package and its config; /var/lib/apt/lists is the package index apt-get update downloads.)
<missing> RUN /bin/sh -c rm -rf /var/lib/apt/lists/* # b… 0B
<missing> RUN /bin/sh -c apt-get purge -y build-essent… 1.2kB
<missing> RUN /bin/sh -c apt-get update && apt-get ins… 317MB
The image is 317MB bigger, forever. The fix is to create and delete in the same RUN, so the temporary files never reach a layer:
RUN apt-get update \
&& apt-get install -y --no-install-recommends curl ca-certificates \
&& rm -rf /var/lib/apt/lists/*
--no-install-recommends matters too: Ubuntu installs "recommended" extra packages by default, often doubling what you install. The Debian and Ubuntu base images already delete downloaded .deb files automatically (/etc/apt/apt.conf.d/docker-clean), so apt-get clean adds nothing; the package lists are what you must remove.
The second rule: changing metadata copies the file
overlayfs cannot change a file's owner or mode in a lower layer. It copies the whole file up into the new layer (a copy-up) and changes it there:
COPY target/app.jar /app/app.jar # +22MB
RUN chown -R 1000:1000 /app # +22MB again
<missing> RUN /bin/sh -c chown -R 1000:1000 /app # b… 22MB
<missing> COPY target/app.jar /app/app.jar # buildkit 22MB
Your JAR is now stored twice. With a 400MB folder of dependencies it is 400MB twice. Set ownership as you copy:
COPY --chown=1000:1000 target/app.jar /app/app.jar
(and --chmod=755 exists for modes, Ch 4). Same result, one copy.
The third rule: COPY'd files belong to root
Whatever USER says, COPY and ADD create files owned by root unless you pass --chown. So this common Dockerfile
USER 1000
COPY . /app
produces an app that cannot write to its own directory:
# app:1 = an image built from that Dockerfile
docker run --rm app:1 touch /app/cache.db
touch: cannot touch '/app/cache.db': Permission denied
USER changes who runs things; it does not change who owns things copied in.
What to look for, in order
When an image is too big, run docker history --no-trunc and look for:
- A huge base at the bottom - a JDK,
node:22,python:3.12(the full variants are around 1GB). Switch to a smaller runtime base (later lessons). - A huge COPY - usually
COPY . .with no.dockerignore:.git,target/,node_modules/(Node's dependency folder). - A RUN that downloads (package managers, dependency caches) whose cleanup happens in a different RUN.
- The same size twice - a
chown -Rorchmod -Rafter a COPY. - Build tools in the final image - compilers, Maven,
-devpackages.
The history tells you which row; the Dockerfile tells you why.
What you can now do
- Read
docker historyand name the row that makes an image big. - Explain why a cleanup in its own RUN frees nothing.
- Avoid the
chown -Rcopy and the root-owned COPY.