OnCallReady

Lesson 10.13 · Images & Builds · 19 min read

The build cache, and why instruction order decides build time

In plain words

Imagine you do the same homework sheet every day, and you keep yesterday's answers. As long as each question is exactly the same as yesterday, you copy the answer. But the moment one question is different, you have to redo that one and every question after it, because the later ones build on it.

That is the build cache. Each instruction has a key: the previous layer plus the command text, or plus a checksum of the files for COPY. If the key matches, BuildKit prints CACHED and skips it; the first miss makes everything below rebuild. So you put what rarely changes (the base, COPY pom.xml and the dependency download) above what changes on every commit (COPY src).

Why this matters

A build that takes 72 seconds on every commit, for every developer and every automated build, adds up to hours a day. Usually 56 of those seconds are re-downloading the same dependencies. BuildKit can skip any step it has run before - if you order the Dockerfile so it can.

What you need to know already: the Dockerfile and reading build output (two lessons ago), docker history (previous lesson), the orders app and Maven words from the Dockerfile lesson.

How the cache decides

BuildKit remembers the result of every step it ran: the build cache. Before running a step it computes the step's cache key - a fingerprint of everything that step depends on. If an earlier build had the same key, it reuses the result and prints CACHED (a cache hit); otherwise it runs the step (a miss). The key is:

FROM         the resolved base image (by digest)
RUN          the parent layer + the exact command string + build args in scope
COPY / ADD   the parent layer + a checksum of the CONTENT of the files copied
metadata     the parent + the instruction text

"The parent layer" is the key of the step before. Three consequences:

1. A miss is contagious. Once one step rebuilds, every step after it rebuilds, whether or not it would have hit. The cache is a chain.

2. COPY hashes content, not timestamps. touch changes the file's modification time (mtime) and nothing else, so BuildKit still says CACHED. Changing one byte of one file does bust (invalidate) it:

$ cd ~/oncall-lab/labs/1c-docker/orders
$ touch src/main/java/com/ing/orders/OrdersApplication.java && docker build -t orders .
 => CACHED [3/4] COPY . .                                             0.0s
 => CACHED [4/4] RUN mvn package -DskipTests                          0.0s
$ echo "// fix" >> src/main/java/com/ing/orders/OrdersApplication.java && docker build -t orders .
 => [3/4] COPY . .                                                    0.1s
 => [4/4] RUN mvn package -DskipTests                                71.7s

(The legacy builder behaved the same; people "remember" mtime mattering because their editors or git checkouts changed the content, or because .git was in the context.)

3. RUN is keyed by its text, not by what it does. RUN apt-get update is cached forever once it succeeds - it does not know the package mirror changed. That is the classic stale-cache bug:

RUN apt-get update                       # cached from last month
RUN apt-get install -y curl=8.5.0-2ubuntu10.6   # new line: runs, finds month-old lists, 404

Always update and install in the same RUN, so changing the package list re-runs the update too.

The ordering rule

Put what changes rarely before what changes constantly. Source code changes on every commit; dependencies change weekly; the base image monthly.

# WRONG: every source edit re-resolves every dependency
FROM maven:3.9-eclipse-temurin-21
WORKDIR /app
COPY . .
RUN mvn package -DskipTests
# RIGHT: the dependency layer survives source edits
FROM maven:3.9-eclipse-temurin-21
WORKDIR /app
COPY pom.xml .
RUN mvn dependency:go-offline
COPY src ./src
RUN mvn package -DskipTests

pom.xml is the file that lists the dependencies; mvn dependency:go-offline downloads them all without building anything. After a source change, the right version rebuilds only the last two steps:

 => CACHED [2/6] WORKDIR /app                                         0.0s
 => CACHED [3/6] COPY pom.xml .                                       0.0s
 => CACHED [4/6] RUN mvn dependency:go-offline                        0.0s
 => [5/6] COPY src ./src                                              0.1s
 => [6/6] RUN mvn package -DskipTests                                15.7s

From 72 seconds to 16, on every commit. The same shape exists for every language: copy the file that lists the dependencies (and its lockfile, which records the exact version of each one), install, then copy the source.

Maven    COPY pom.xml .                  RUN mvn dependency:go-offline   COPY src ./src
Gradle   COPY build.gradle settings.gradle ./   RUN gradle dependencies   COPY src ./src
Node     COPY package.json package-lock.json ./   RUN npm ci   COPY . .
Python   COPY requirements.txt .         RUN pip install -r requirements.txt   COPY . .
Go       COPY go.mod go.sum ./           RUN go mod download   COPY . .

(Gradle is another Java build tool; npm installs Node.js packages; pip installs Python packages; go mod download fetches Go modules.)

mvn dependency:go-offline does not always fetch every plugin a later package needs; the first build after a pom change may still download a few things. It gets you 95% of the win, which is why it is the standard idiom.

ARG and ENV bust the cache too

Changing an ENV or ARG value changes the key of that step, and the miss spreads from there. An ARG is subtler: it does not bust anything where it is declared, but every RUN after it includes the value in its key:

FROM node:22-slim
ARG BUILD_DATE           # the build server passes --build-arg BUILD_DATE=$(date +%s)
WORKDIR /app
COPY package*.json ./
RUN npm ci               # key includes BUILD_DATE: misses on EVERY build
COPY . .

(date +%s prints the current time in seconds, so the value is new every build.) A value that changes every build (a date, the git commit id, a build number), declared near the top, silently disables the cache for everything below it. Declare such args as late as possible, just before the one step that uses them - usually a LABEL at the very end:

COPY . .
ARG GIT_SHA
LABEL org.opencontainers.image.revision=$GIT_SHA

Cache mounts: keep the download cache across misses

Even with perfect ordering, the dependency step reruns when the dependency list changes - and re-downloads the world. A cache mount gives one RUN a directory that persists between builds but is not part of any layer:

RUN --mount=type=cache,target=/root/.m2 mvn package -DskipTests
RUN --mount=type=cache,target=/root/.npm npm ci
RUN --mount=type=cache,target=/root/.cache/pip pip install -r requirements.txt
RUN --mount=type=cache,target=/go/pkg/mod go build -o /out/app .

target= is the directory where the tool keeps its downloads. Two wins at once: the layer is smaller (the cache directory is not in it), and a miss only downloads what changed. The cache lives in the builder on this machine, so it helps on your laptop and on a build server that keeps its disk - not on a fresh build machine each time. The cache-mounts lesson later in this chapter covers that case.

Forcing things

docker build --no-cache .        ignore the cache for every step
docker build --pull .            re-check FROM against the registry (new base digest)
docker builder prune             drop the build cache (it can grow to many GB)
docker system df                 how big the build cache is right now

--pull is the one people forget. Without it, BuildKit uses the base image you already have locally, so your "fresh" build can sit on a base that is months old. Automated builds should always --pull.

Reading a build for cache problems

When a build is slow, read the step list top to bottom and find the first step that is not CACHED. Everything below it is collateral. Then ask what changed in that step's key: a file it copies, its command text, an ARG above it, the base image. That one question solves almost every "why does this layer keep re-running".

What you can now do

Why it helps

Every developer and every CI run pays for your Dockerfile ordering, on every commit. The lesson's example goes from 72 seconds to 16 by copying pom.xml first; on a team of ten pushing all day, that is hours. The debugging skill is concrete: read the step list, find the first step that is not CACHED, and ask what changed in its key. That is how you find the ARG BUILD_DATE at the top of a teammate's Dockerfile that silently disabled the cache for every build. It also explains a real incident: RUN apt-get update on its own line, cached for a month, then a new install line fails with 404s. And in CI you now know why --pull matters: without it, the 'fresh' image may sit on a months-old base.

Commands in this lesson

cd touch echo

FAQ

Does touching a file bust the COPY cache?

No. BuildKit hashes file content, not timestamps, so touch alone still gives CACHED. Changing one byte does bust it. The legacy builder also ignored mtimes. People remember otherwise because editors, formatters or git checkouts actually changed content, or because .git was in the context and changes on every commit.

Why does RUN apt-get update never refresh?

A RUN step's cache key is the parent layer plus the exact command text. BuildKit does not know the mirror changed, so once RUN apt-get update succeeded it stays cached forever. A new install line below it then runs against month-old package lists and fails with 404s. Always run update and install in the same RUN, so changing the package list reruns the update.

How can an ARG break caching if it does not create a layer?

Declaring an ARG does nothing to the cache by itself. But RUN steps after it include the ARG's value in their key. If CI passes a value that changes every build, a date, the git SHA or a build number, and the ARG is declared near the top, every RUN below it misses every time. Declare such args just before the step that needs them, usually a LABEL at the end.

What is the difference between --no-cache and --pull?

--no-cache ignores the build cache for every step and rebuilds everything, which is slow and rarely what you want. --pull only re-resolves the FROM images against the registry, so you get the current base digest; if it did not change, the cache still hits. CI builds should use --pull so a new base patch is picked up; --no-cache is a debugging tool.

Does mvn dependency:go-offline really make the build work offline?

Not fully. It downloads most dependencies and plugins, but some plugins that package needs later are resolved only at that point, so the first build after a pom change may still download a few things. It still gets most of the win, which is why it is the standard idiom. A cache mount on /root/.m2 covers the rest.

In an interview Junior

How would you order a Dockerfile for a Java or Node service to use the build cache well?

BuildKit reuses a step (CACHED) when its cache key matches: the parent step plus the command text, or the content of the files a COPY copies. A miss is contagious: every step after it rebuilds.

So put what changes rarely before what changes constantly:

COPY pom.xml .
RUN mvn dependency:go-offline
COPY src ./src
RUN mvn package -DskipTests

For Node: COPY package.json package-lock.json ./, RUN npm ci, then COPY . .. Source edits now skip the dependency download.

Also: apt-get update and install in the same RUN (a RUN is keyed by its text, so a lone update stays cached forever); declare an ARG whose value changes every build (date, commit SHA) as late as possible, because every RUN after it includes it in its key; and use --pull in automated builds so the base is current. When a build is slow, find the first step that is not CACHED.

Also asked: Why would a build that took 20 seconds suddenly take 4 minutes on every run? · What does docker build --no-cache do, and what does --pull do? · Why should apt-get update and apt-get install be in the same RUN?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.