OnCallReady

Lesson 10.19 · Images & Builds · 15 min read

Cache mounts, remote caches, and builds in CI

In plain words

There are two ways to save time on homework. One: if the question is exactly the same as yesterday, copy yesterday's answer and skip it. Two: if the question changed and you must redo it, at least keep your dictionary on the desk instead of walking to the library every time.

The layer cache is the first kind: a step with the same inputs is skipped entirely (CACHED). A cache mount, RUN --mount=type=cache,target=/root/.m2, is the dictionary: a directory that survives between builds but is never part of any layer, so a step that must rerun finds its downloads already there. Both live in the builder, which is why a fresh CI runner has neither unless you export the cache to a registry.

The problem

You ordered the Dockerfile well: pom.xml first, dependencies, then source. A code change now skips the dependency step. But when someone adds a dependency, pom.xml changes, the step must run - and Maven downloads every library again, from zero, because the last download is gone. Eighty seconds, several times a week, for everyone.

This lesson adds the second kind of cache that fixes that, and then looks at what happens to both caches on a CI server.

What you need to know already: the build cache and "the first miss rebuilds everything after it" (10.13 and its missions), pom.xml and mvn dependency:go-offline (10.14), what CI is (10.18). A stage is one FROM section of a Dockerfile; AS build names it (10.18 showed one; the multi-stage lesson explains them).

Two different caches

BuildKit has two caches, and people mix them up:

layer cache    "this exact step with these exact inputs was built before"
               -> the whole step is skipped (CACHED). All or nothing.
cache mount    a directory that survives between builds but is not in any layer
               -> the step still RUNS, but finds its downloads already there.

The layer cache is what the ordering rules optimise. It cannot help when a step's input really changed.

A cache mount is a folder that BuildKit keeps on the build machine and plugs into one RUN step while it runs - like a USB stick that is inserted for the step and removed before the layer is saved. You ask for one with RUN --mount=type=cache,target=<dir>:

# syntax=docker/dockerfile:1
FROM maven:3.9-eclipse-temurin-21 AS build
WORKDIR /app
COPY pom.xml .
COPY src ./src
RUN --mount=type=cache,target=/root/.m2 mvn package -DskipTests

The first line, # syntax=docker/dockerfile:1, tells BuildKit to use the current Dockerfile syntax, which is where --mount lives.

first build           => [5/5] RUN --mount=type=cache,target=/root/.m2 mvn package -DskipTests   80.7s
pom.xml changed       => [5/5] RUN --mount=type=cache,target=/root/.m2 mvn package -DskipTests   15.7s

The step reran (its input changed), but Maven found /root/.m2 already full and downloaded only the new dependency. And the image got smaller, because the repository is not in the layer:

# cm:1 = the image you build in the mission below
docker history cm:1
IMAGE          CREATED          CREATED BY                                      SIZE      COMMENT
8a9a210c3256   20 seconds ago   RUN /bin/sh -c mvn package -DskipTests # bui…   24.5MB    buildkit.dockerfile.v0

Read it as usual: one row per instruction, SIZE = what that step added. Without the mount that row would be ~180MB: the jar plus the whole .m2.

The targets per ecosystem

Each build tool keeps its downloads in a known folder; that folder is the target:

Maven     --mount=type=cache,target=/root/.m2
Gradle    --mount=type=cache,target=/root/.gradle
npm       --mount=type=cache,target=/root/.npm
pip       --mount=type=cache,target=/root/.cache/pip
Go        --mount=type=cache,target=/go/pkg/mod  --mount=type=cache,target=/root/.cache/go-build
apt       --mount=type=cache,target=/var/cache/apt,sharing=locked

Two gotchas:

Where cache mounts live

Cache mounts are builder-local: they live on the machine where BuildKit runs, next to the layer cache. Two commands to look at and clean them:

$ docker buildx du
ID                         RECLAIMABLE   SIZE      LAST ACCESSED
m2o0qb3y1q4x...            true          182MB     2 minutes ago
...
Reclaimable:   1.83GB
Total:         1.83GB
$ docker builder prune -af

CI: the runner forgets everything

A hosted CI runner (a build machine the CI service provides, such as GitHub's) is a fresh virtual machine every job: no layer cache, no cache mounts, no base images. Every build is a cold build unless you save the cache somewhere and load it back next time. The usual place is the registry itself:

docker buildx build \
  --cache-from type=registry,ref=registry.lab/team/orders:buildcache \
  --cache-to   type=registry,ref=registry.lab/team/orders:buildcache,mode=max \
  -t registry.lab/team/orders:3f9c2ab --push .

GitHub's CI also has its own store, type=gha. Cache mounts are not saved this way - on fresh runners the layer ordering does most of the work, which is why getting it right matters so much.

A self-hosted runner (your own server with Docker installed, kept between jobs) keeps both caches, which is fast and comes with the usual caveats: the disk fills (prune on a schedule - a systemd timer from Ch 2 does it) and builds from different branches share one cache.

--pull in CI, always

The layer cache is keyed on the base image digest. Without --pull, BuildKit uses whatever base the runner has locally - on a self-hosted runner, possibly one from months ago, missing every security fix since. --pull checks the registry for the current digest of the FROM tag; if the base did not change, the cache still hits.

A CI build line worth copying

docker build --check .
docker build --pull --progress=plain \
  --build-arg GIT_SHA="$GIT_SHA" \
  -t "registry.lab/team/orders:$VERSION" \
  -t "registry.lab/team/orders:$GIT_SHA" .
docker push --all-tags registry.lab/team/orders

--check to lint first, plain progress for readable logs, two tags that never move (the version and the commit SHA), the SHA passed as a build arg declared at the end of the Dockerfile (so it does not bust the cache - the ARG trap from 10.15), and --all-tags to push both.

Later (Ch 25): the same four lines become a pipeline definition that the CI system runs on every push.

What you can now do

Why it helps

The hosted CI runner is a fresh VM for every job: no layers, no cache mounts, no base images. That is why a build that takes 20 seconds on your laptop takes 6 minutes in the pipeline, and it is one of the most common platform tickets. Knowing the two caches lets you answer it properly: ordering fixes the layer cache, --cache-to type=registry,mode=max carries it between jobs, and cache mounts help on persistent runners and laptops. A cache mount also makes the image smaller, since .m2 never lands in a layer. And you can explain the caveats of self-hosted runners with a persistent Docker: fast, but the disk fills up, and different branches share one cache, so docker buildx du and scheduled pruning belong in the runbook.

Commands in this lesson

docker

FAQ

What is the difference between the layer cache and a cache mount?

The layer cache skips a whole step when its inputs match a previous build: all or nothing. A cache mount does not skip anything; the step still runs, but a directory like /root/.m2 or /root/.npm persists between builds, so the package manager only downloads what is new. Layer cache for unchanged steps, cache mount for steps that must rerun.

Are cache mounts included in the image?

No. The mounted directory exists only while that RUN executes and is stored in the builder, not in any layer. That is why the image gets smaller: the Maven repository or npm cache is not baked in. It also means anything your app needs at runtime must be copied out of the cache directory into a real path during the RUN.

Why do cache mounts not help on GitHub's hosted runners?

They live in the builder's local storage, and a hosted runner is a fresh VM each job, so the cache starts empty. Registry cache export (--cache-to / --cache-from) or type=gha carries layer cache between jobs, but cache mounts are not exported that way. On ephemeral runners, good layer ordering does most of the work.

What does mode=max do in --cache-to?

By default (mode=min) only the layers of the final image are exported. mode=max exports the layers of every stage, including intermediate build stages. For multi-stage builds that is what you want, because the expensive dependency step lives in the build stage, which is not part of the final image. It costs more registry storage.

Can I use a cache mount together with pip --no-cache-dir?

They work against each other. --no-cache-dir tells pip not to keep its download cache, so the cache mount stays empty and helps nothing. Choose one: --no-cache-dir to keep a plain layer small, or a cache mount on /root/.cache/pip without the flag to keep the image small and reuse downloads.

In an interview Junior

Why are Docker builds slower in CI than on a laptop, and what can you do about it?

A hosted CI runner is a fresh virtual machine for every job: no layer cache, no cache mounts, no base images. Every build is cold.

What helps:

On a machine that keeps its disk, a cache mount (RUN --mount=type=cache,target=/root/.m2 mvn package) keeps Maven's downloads between builds, so a changed pom.xml downloads only what is new - and they stay out of the layer. Cache mounts are builder-local; the registry cache does not carry them.

Also asked: What is the difference between the layer cache and a cache mount? · How do you see and clean the build cache? · Why should CI builds always use --pull?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.