OnCallReady

Lesson 10.24 · Images & Builds · 17 min read

Multi-stage builds for Java: from 1GB to under 200MB

In plain words

To build a wooden chair you need a workshop: saws, glue, clamps, sawdust everywhere. To use the chair you need only the chair. Nobody delivers the whole workshop to your living room along with it.

A multi-stage Dockerfile has a workshop stage (FROM maven:... AS build) with the JDK, Maven and hundreds of MB of dependencies, and a final stage that starts fresh from a small runtime base and takes only the finished jar with COPY --from=build. Only the last stage becomes the image. jlink goes one step further and builds a Java runtime with only the modules your app uses, and layered jars split the jar so that a code change ships a 660kB layer instead of 22MB.

Why this matters

To build a Java service you need a JDK, Maven and a few hundred MB of downloaded dependencies. To run it you need a Java runtime (JRE) and one JAR. An image built in one piece ships both:

$ docker images orders
REPOSITORY   TAG     IMAGE ID       CREATED          SIZE
orders       naive   cf3ab9a158c8   2 minutes ago    945MB

A compiler, a build tool, your source code and ~/.m2 (Maven's download folder) - downloaded by every server that runs the app, flagged by every security scanner, and available to anyone who breaks in. A multi-stage build keeps the workshop out of the delivery.

What you need to know already: the Dockerfile instructions and the Java words JAR / JDK / JRE / Maven ("The Dockerfile, and reading a build"), docker history, cache ordering (the cache lesson).

Two stages

FROM maven:3.9-eclipse-temurin-21 AS build
WORKDIR /app
COPY pom.xml .
RUN mvn dependency:go-offline
COPY src ./src
RUN mvn package -DskipTests

FROM eclipse-temurin:21-jre
WORKDIR /app
COPY --from=build /app/target/orders-2.14.1.jar app.jar
USER 1000:1000
ENTRYPOINT ["java","-jar","app.jar"]
$ docker images orders
REPOSITORY   TAG     IMAGE ID       CREATED          SIZE
orders       slim    9d2e3f4a5b6c   10 seconds ago   293MB
orders       naive   cf3ab9a158c8   5 minutes ago    945MB
# orders:slim = the multi-stage image from the mission below
docker history orders:slim
IMAGE          CREATED          CREATED BY                                      SIZE      COMMENT
9d2e3f4a5b6c   10 seconds ago   ENTRYPOINT ["java" "-jar" "app.jar"]            0B        buildkit.dockerfile.v0
<missing>      10 seconds ago   USER 1000:1000                                  0B        buildkit.dockerfile.v0
<missing>      10 seconds ago   COPY /app/target/orders-2.14.1.jar app.jar #…   22MB      buildkit.dockerfile.v0
<missing>      10 seconds ago   WORKDIR /app                                    0B        buildkit.dockerfile.v0
<missing>      3 weeks ago      ENTRYPOINT ["/__cacert_entrypoint.sh"]          0B        buildkit.dockerfile.v0
...

The top four rows are yours: 22MB. Everything below is eclipse-temurin:21-jre (271MB). No Maven row, no JDK row.

Why not just "use a smaller base"?

The runtime base is now most of the image. The options, with real sizes:

eclipse-temurin:21-jdk             489MB   full JDK: javac, jlink, jshell - build only
eclipse-temurin:21-jre             271MB   Ubuntu + JRE. The safe default
gcr.io/distroless/java21-debian12  226MB   Debian libs + JRE, no shell, no package manager
eclipse-temurin:21-jre-alpine      168MB   musl libc - read the base-image lesson first
debian:12-slim + jlink runtime     ~115MB  a JRE with only the modules you use

(javac is the Java compiler. distroless images and alpine/musl get their own lesson soon; for now: distroless = only the runtime, no shell; alpine = a tiny Linux with a different C library.)

With a normal base, the stock JRE alone puts you over 200MB. The plan's target (under 200MB, multi-stage, non-root) needs one more move: build your own Java runtime.

jlink: a JRE with only the modules you use

Java's standard library is split into modules (java.base, java.logging, java.sql...). The JDK ships jlink, a tool that assembles a runtime from just the modules you list. A typical Java web service needs about eight:

FROM eclipse-temurin:21-jdk AS jre
RUN jlink \
      --add-modules java.base,java.logging,java.naming,java.management,java.security.jgss,java.instrument,java.sql,jdk.unsupported \
      --strip-debug --no-man-pages --no-header-files --compress=zip-6 \
      --output /javaruntime

FROM debian:12-slim
ENV JAVA_HOME=/opt/java/openjdk
ENV PATH="${JAVA_HOME}/bin:${PATH}"
COPY --from=jre /javaruntime $JAVA_HOME

The flags: --add-modules is the list; --strip-debug, --no-man-pages and --no-header-files leave out things only developers need; --compress=zip-6 compresses the result; --output is where to write it. JAVA_HOME and PATH tell the shell and tools where the new java lives.

<missing>   COPY /javaruntime $JAVA_HOME # buildkit      38.2MB
<missing>   # debian.sh --arch 'arm64' out/ 'bookworm' …  74.8MB

75MB of Debian + 38MB of Java + 22MB of JAR = about 135MB. jdeps --print-module-deps --ignore-missing-deps app.jar reads your JAR and prints the module list it needs. Get the list wrong and you find out only when the app runs (java.lang.NoClassDefFoundError - Java could not find a piece of code), so test the image, do not just build it. (--compress=zip-6 is the JDK 21 spelling; older guides use --compress=2.)

Three stages, and BuildKit only builds what it needs

FROM maven:3.9-eclipse-temurin-21 AS build
...
FROM eclipse-temurin:21-jdk AS jre
...
FROM debian:12-slim AS runtime
COPY --from=jre /javaruntime /opt/java/openjdk
COPY --from=build /app/target/orders-2.14.1.jar /app/app.jar

FROM build AS test
RUN mvn verify

(mvn verify runs the tests.) The build and jre stages do not depend on each other, so BuildKit builds them in parallel. And a stage the target does not need is skipped entirely: docker build --target runtime . never runs the test stage, docker build --target test . never runs jre. The legacy builder ran every stage in order - another reason to have buildx installed.

Layered jars: stop re-shipping 20MB for a one-line change

orders is a Java web app, built on a popular framework for web services. It is packaged as a fat jar: your code and every dependency in one file. Change one line and the whole JAR is a new 22MB layer: every deploy uploads and downloads 22MB. The framework can unpack the JAR into folders that change at different rates, and you copy each as its own layer:

FROM maven:3.9-eclipse-temurin-21 AS build
WORKDIR /app
COPY pom.xml .
RUN mvn dependency:go-offline
COPY src ./src
RUN mvn package -DskipTests
RUN java -Djarmode=tools -jar target/orders-2.14.1.jar extract --layers --launcher --destination extracted

FROM eclipse-temurin:21-jre
WORKDIR /app
COPY --from=build /app/extracted/dependencies/ ./
COPY --from=build /app/extracted/spring-boot-loader/ ./
COPY --from=build /app/extracted/snapshot-dependencies/ ./
COPY --from=build /app/extracted/application/ ./
ENTRYPOINT ["java","org.springframework.boot.loader.launch.JarLauncher"]
<missing>   COPY /app/extracted/application/ ./ # buildkit            660kB
<missing>   COPY /app/extracted/snapshot-dependencies/ ./ # bu…      0B
<missing>   COPY /app/extracted/spring-boot-loader/ ./ # build…      300kB
<missing>   COPY /app/extracted/dependencies/ ./ # buildkit          20.5MB

The dependencies layer only changes when pom.xml does. A code change is now a 660kB layer; the registry and every server already have the other 20MB. -Djarmode=tools with extract --layers --launcher is the form in current versions (3.3+); older projects use -Djarmode=layertools -jar app.jar extract and the launcher class org.springframework.boot.loader.JarLauncher.

Later (Ch 21): the framework is Spring Boot; Ch 21 covers how it starts, serves and shuts down.

The four moves, in order of impact

For the interview question "take this 1.2GB image to 200MB":

  1. Multi-stage. The build toolchain never reaches the final image. The biggest single win, and it removes a compiler from production.
  2. A runtime base, not a build base. JRE instead of JDK or Maven; slim instead of full; jlink or distroless for Java.
  3. A .dockerignore, so COPY cannot drag in .git, target/ or node_modules.
  4. Clean up in the same RUN you created the mess in, and never chown -R after a COPY.

Then prove it with docker history, not with docker images: the history says which row is still too big.

What you can now do

Why it helps

'Take this 1.2GB image to under 200MB' is one of the most common practical interview tasks for platform roles, and it is also a real ticket: every server pulls that image when it is added, every scanner reports every CVE in Maven and the JDK, and anyone who gets a shell in production finds a compiler. With this lesson you can do it in order of impact and prove each step with docker history. In reviews you will recognise a single-stage Java or Node image and know what to suggest. And it explains why BuildKit matters beyond speed: it builds independent stages in parallel and skips stages the target does not need, so a test stage costs nothing when you build --target runtime.

Commands in this lesson

docker

FAQ

Which stage becomes the image?

The last stage in the Dockerfile, unless you pick another with --target <name>. Earlier stages exist only during the build; nothing from them reaches the image unless a later stage explicitly copies it with COPY --from=<stage>. You can also copy from an external image, like COPY --from=nginx:1.27 /etc/nginx/nginx.conf ..

Why not just use the JDK image as the runtime?

Because the JDK includes the compiler, jlink, jshell and more, none of which a running service needs. eclipse-temurin:21-jdk is about 489MB, the JRE about 271MB. Beyond size, a compiler in production helps an attacker. The JRE is the safe default; jlink or distroless are for when you need to go smaller.

What happens if my jlink module list is wrong?

The image builds fine, and the app fails at runtime, typically with java.lang.NoClassDefFoundError or a missing service provider when a code path first touches a module you left out. That is why you derive the list with jdeps --print-module-deps --ignore-missing-deps and then actually run and exercise the image, not just build it.

Does BuildKit run stages that the final image does not need?

No. BuildKit builds a dependency graph and only builds the stages the target needs, in parallel where they are independent. So docker build --target runtime skips a test stage entirely. The legacy builder ran every stage in order, whether needed or not.

What do layered jars actually buy me?

Smaller deploys. A fat jar bundles your code with every dependency, so a one-line change is a new 22MB layer to push and pull. Extracting it with -Djarmode=tools ... extract --layers --launcher into dependencies, loader, snapshot dependencies and application lets you copy each as its own layer. Dependencies change only with the pom, so a code change becomes a layer of a few hundred kB.

In an interview Junior

What is a multi-stage Docker build and why would you use it?

A Dockerfile with several FROM lines. Each FROM starts a new stage with a fresh filesystem; AS build names it, and only the last stage (or the one picked with --target) becomes the image. COPY --from=build takes exactly the files you name from another stage.

FROM maven:3.9-eclipse-temurin-21 AS build
WORKDIR /app
COPY pom.xml .
RUN mvn dependency:go-offline
COPY src ./src
RUN mvn package -DskipTests

FROM eclipse-temurin:21-jre
COPY --from=build /app/target/app.jar /app/app.jar
ENTRYPOINT ["java","-jar","/app/app.jar"]

Why: you need a JDK, Maven and hundreds of MB of dependencies to build, but only a JRE and one JAR to run. The build toolchain never reaches production: a much smaller image, fewer packages for scanners to flag, no compiler for an attacker. BuildKit also builds independent stages in parallel and skips stages the target does not need.

Also asked: How would you get a Java image under 200MB? · What does jlink do? · What does docker build --target do?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.