OnCallReady

Images & Builds: interview questions

The question you are most likely to get for each topic, a model answer, and what else comes up. From chapter 10 of the course.

How would you take a 1.2GB Java image down to under 200MB? Junior

First docker history --no-trunc to see which row is big, then the four moves in order of impact:

  1. Multi-stage build - build with maven/JDK in a build stage, COPY --from=build only the JAR into the final stage. Maven, the JDK, ~/.m2 and the source stay behind.
  2. A runtime base - eclipse-temurin:21-jre instead of a JDK or Maven image; for under 200MB, a slim Debian base plus a runtime built with jlink (only the modules the app uses).
  3. A .dockerignore so COPY cannot drag in .git, target/ or node_modules.
  4. Clean up in the same RUN you created the mess in (a later RUN only adds a whiteout), and COPY --chown instead of chown -R.

Then prove it with docker history again, and test that the app still starts - a missing jlink module only fails at runtime.

Also asked: What is the difference between an image and a container? · Why should you not use the latest tag in production? · Why should a container not run as root?

docker ps fails. How do you tell whether it is a permissions problem or the daemon being down? Junior

The docker command is only a client; it talks to dockerd over the Unix socket /run/docker.sock (owner root, group docker, mode 660). Read the error:

docker version shows both: a Server: section only when the daemon answers. And remember the docker group is equivalent to root.

Also asked: Why is membership of the docker group equivalent to root? · What is the difference between the docker client and dockerd? · What is the risk of mounting /var/run/docker.sock into a container?

Learn it: 10.1 Installing Docker, the daemon, and the socket

Explain what a container is without saying "lightweight VM". Junior

A container is a normal Linux process on the host kernel, isolated with namespaces and limited with cgroups.

There is no hypervisor and no guest kernel. So it starts in milliseconds, it is visible in the host's ps, uname -r shows the host kernel, an amd64 binary fails on arm64 with exec format error, and a kernel exploit is not contained.

Also asked: What is the difference between an image and a container? · Why does a file written inside a container disappear when you remove it? · Why does free inside a container show the host's memory?

Learn it: 10.3 What a container actually is

Why is using the latest tag in production a bad idea? Junior

A tag is a movable pointer in the registry, and latest is not special: it is just the tag you get when you did not type one. It does not mean "newest".

What goes wrong:

Instead: deploy tags that never move (2.14.1, or the git commit id), and pin the digest (name@sha256:...) for anything critical - the digest is the hash of the image's manifest, so it can only ever mean those exact bytes. docker images --digests or docker inspect -f '{{index .RepoDigests 0}}' show it. Pin base images by digest too, and let a bot like Renovate propose updates.

Also asked: What is the difference between a tag and a digest? · What are the parts of an image reference like registry.lab/team/orders:2.14.1? · Does the CREATED column of docker images show when you pulled the image?

Learn it: 10.5 Images, tags and digests

Which Dockerfile instructions create layers, and why does that matter? Junior

Why it matters: layers are what the image weighs and what every server downloads, and in BuildKit output only filesystem steps are numbered ([3/4]), with the time each took on the right - that is where build time goes. Metadata still counts for the build cache.

Two classics to mention: EXPOSE publishes nothing (only docker run -p does), and prefer COPY over ADD, which also unpacks tarballs and downloads URLs. When a build fails, the output shows the failing line with >>> and its exit code (127 = command not found); --progress=plain shows the full output of each step.

Also asked: What is the difference between COPY and ADD? · Does EXPOSE in a Dockerfile open a port on the host? · How do you read the output of a failed docker build?

Learn it: 10.8 The Dockerfile, and reading a build

How do you find out why a Docker image is so large? Junior

docker history --no-trunc IMAGE lists every instruction with the bytes it added, newest at the top; the base image is the bottom rows. Look for, in order:

  1. A huge base - a JDK, node:22, a full python image.
  2. A huge COPY - usually COPY . . without a .dockerignore (.git, target/, node_modules/).
  3. A RUN that downloads whose cleanup is in a different RUN. Deleting in a later layer frees nothing: it only adds a whiteout; the bytes stay in the earlier layer. Create and delete in the same RUN (apt-get update && apt-get install -y --no-install-recommends ... && rm -rf /var/lib/apt/lists/*).
  4. The same size twice - chown -R after a COPY makes a copy-up of every file; use COPY --chown.
  5. Build tools in the final image - compilers, Maven.

The history says which row; the Dockerfile says why.

Also asked: Why does RUN rm -rf on its own line not make the image smaller? · Why are files you COPY owned by root even after USER? · How would you check whether a password was written into an image's history?

Learn it: 10.10 Reading docker history: where the bytes went

How would you order a Dockerfile for a Java or Node service to use the build cache well? Junior

BuildKit reuses a step (CACHED) when its cache key matches: the parent step plus the command text, or the content of the files a COPY copies. A miss is contagious: every step after it rebuilds.

So put what changes rarely before what changes constantly:

COPY pom.xml .
RUN mvn dependency:go-offline
COPY src ./src
RUN mvn package -DskipTests

For Node: COPY package.json package-lock.json ./, RUN npm ci, then COPY . .. Source edits now skip the dependency download.

Also: apt-get update and install in the same RUN (a RUN is keyed by its text, so a lone update stays cached forever); declare an ARG whose value changes every build (date, commit SHA) as late as possible, because every RUN after it includes it in its key; and use --pull in automated builds so the base is current. When a build is slow, find the first step that is not CACHED.

Also asked: Why would a build that took 20 seconds suddenly take 4 minutes on every run? · What does docker build --no-cache do, and what does --pull do? · Why should apt-get update and apt-get install be in the same RUN?

Learn it: 10.13 The build cache, and why instruction order decides build time

Why are Docker builds slower in CI than on a laptop, and what can you do about it? Junior

A hosted CI runner is a fresh virtual machine for every job: no layer cache, no cache mounts, no base images. Every build is cold.

What helps:

On a machine that keeps its disk, a cache mount (RUN --mount=type=cache,target=/root/.m2 mvn package) keeps Maven's downloads between builds, so a changed pom.xml downloads only what is new - and they stay out of the layer. Cache mounts are builder-local; the registry cache does not carry them.

Also asked: What is the difference between the layer cache and a cache mount? · How do you see and clean the build cache? · Why should CI builds always use --pull?

Learn it: 10.19 Cache mounts, remote caches, and builds in CI

What is the Docker build context, and why does .dockerignore matter? Junior

The last argument of docker build (the .) is the build context: a directory sent to the builder in full before any step runs. COPY can only see files inside it. The output shows its size in transferring context:.

Without a .dockerignore, that is .git, target/, node_modules/, editor settings - and .env with credentials. Slow every build, and COPY . . bakes all of it into a layer that anyone who can pull the image can read.

.dockerignore sits at the context root and lists what not to send. The gotcha: patterns are anchored at the root, unlike .gitignore - node_modules only matches ./node_modules; use **/node_modules for nested ones. ! lines are exceptions, the last match wins.

If a COPY fails with failed to calculate checksum ... not found, the file is outside the context or ignored.

Also asked: A COPY works on your machine but fails in CI with "not found". What do you check? · How is .dockerignore different from .gitignore? · Why is COPY . . risky?

Learn it: 10.21 The build context and .dockerignore

What is a multi-stage Docker build and why would you use it? Junior

A Dockerfile with several FROM lines. Each FROM starts a new stage with a fresh filesystem; AS build names it, and only the last stage (or the one picked with --target) becomes the image. COPY --from=build takes exactly the files you name from another stage.

FROM maven:3.9-eclipse-temurin-21 AS build
WORKDIR /app
COPY pom.xml .
RUN mvn dependency:go-offline
COPY src ./src
RUN mvn package -DskipTests

FROM eclipse-temurin:21-jre
COPY --from=build /app/target/app.jar /app/app.jar
ENTRYPOINT ["java","-jar","/app/app.jar"]

Why: you need a JDK, Maven and hundreds of MB of dependencies to build, but only a JRE and one JAR to run. The build toolchain never reaches production: a much smaller image, fewer packages for scanners to flag, no compiler for an attacker. BuildKit also builds independent stages in parallel and skips stages the target does not need.

Also asked: How would you get a Java image under 200MB? · What does jlink do? · What does docker build --target do?

Learn it: 10.24 Multi-stage builds for Java: from 1GB to under 200MB

What is the difference between the exec form and the shell form of ENTRYPOINT, and why does it matter? Junior

docker stop sends SIGTERM to PID 1. sh does not forward it, and PID 1 in a PID namespace ignores signals it has no handler for. So nothing happens for the 10-second grace period, then SIGKILL: exit 137 (128 + 9), requests killed mid-flight. With the exec form the JVM handles SIGTERM, shuts down cleanly and exits 143.

Fixes: exec form; an entrypoint script that ends in exec "$@"; or an init like tini (docker run --init) for programs without a SIGTERM handler. BuildKit warns with JSONArgsRecommended.

And: ENTRYPOINT is the program, CMD its default arguments; arguments after the image name replace CMD.

Also asked: What is the difference between ENTRYPOINT and CMD? · What do exit codes 137 and 143 mean for a container? · Why should an entrypoint script end with exec "$@"?

Learn it: 10.28 ENTRYPOINT, CMD, and who is PID 1

Outline a production Dockerfile for a Node.js API. Junior

Two stages - a fat one that builds, a thin one that runs:

FROM node:22 AS build
WORKDIR /app
COPY package.json package-lock.json ./
RUN npm ci
COPY . .
RUN npm run build

FROM node:22-slim
ENV NODE_ENV=production
WORKDIR /app
COPY package.json package-lock.json ./
RUN npm ci --omit=dev
COPY --from=build /app/dist ./dist
USER node
CMD ["node","dist/server.js"]

The points to say out loud: manifests first for the cache; npm ci (exactly the lockfile) not npm install; --omit=dev in the runtime stage drops devDependencies (TypeScript, test tools); a -slim runtime base; the image's node user; and run node directly in exec form, not npm start (npm as PID 1 does not forward signals). Node has no default SIGTERM handler, so add process.on('SIGTERM', ...) or use --init.

Also asked: How do you keep gcc out of the final image of a Python service that needs it to install a package? · What is the CGO trap in Go container images? · Why does npm ci beat npm install in a Dockerfile?

Learn it: 10.33 Multi-stage for Node, Python and Go

What are the trade-offs between a full, slim, alpine and distroless base image? Junior

Smaller means fewer packages to patch and pull, and fewer tools when something breaks. To debug a shell-less container, work from next to it: docker logs, docker run --network container:api nicolaka/netshoot, docker cp.

Also asked: How would you debug a container that has no shell? · Why can a Java app with native libraries break on alpine? · What is a CVE, and why do smaller images have fewer of them?

Learn it: 10.38 Choosing a base: slim, alpine, distroless, scratch

Why should containers not run as root, and how do you change it? Junior

Without user namespaces, root in the container is uid 0 on the host kernel. Dropped capabilities, seccomp and AppArmor limit it, but if anything breaks out of a root container, it is root on the host.

In the Dockerfile:

RUN groupadd -r app && useradd -r -g app -u 10001 app
COPY --chown=app:app target/app.jar /app/app.jar
USER 10001

Check it: docker run --rm IMAGE id.

Also asked: How do you pass a private registry token to a build without leaking it into the image? · Why does deleting a copied secret in a later RUN not remove it from the image? · Are ARG and ENV values visible in a built image?

Learn it: 10.40 Non-root users, and secrets that leak into images

What does HEALTHCHECK do in a Dockerfile, and what are its pitfalls? Junior

It tells Docker to run a command inside the container on an interval and record the result: exit 0 healthy, exit 1 unhealthy.

HEALTHCHECK --interval=10s --timeout=3s --start-period=20s --retries=3 \
  CMD curl -fsS http://localhost:8080/actuator/health || exit 1

--start-period gives a slow starter time before failures count. The state shows in docker ps (health: starting, healthy, unhealthy); docker inspect -f '{{json .State.Health}}' shows the last results with their output.

Pitfalls:

Also asked: How would you set the HEALTHCHECK options for a service that takes 20 seconds to start? · Does Docker restart a container that becomes unhealthy? · Why should a health check not depend on the database?

Learn it: 10.44 HEALTHCHECK

Here is a Dockerfile with FROM node:latest, COPY . ., RUN npm install and CMD npm start. What would you change? Junior

Go through it in five passes - base, order, layers, runtime, secrets - and say the cost of each problem:

Then ask for evidence: docker history, docker images, time docker stop.

Also asked: How would you enforce Dockerfile quality across many teams? · What does hadolint check? · What makes a good code review comment on a Dockerfile?

Learn it: 10.50 Reviewing a Dockerfile in five minutes

What is inside a container image, technically? Junior

A few blobs, each named by its sha256 digest:

Everything references everything else by digest, so it is tamper-evident. Which ID is which: the image digest you pin is the manifest's digest; the IMAGE ID in docker images is the config's; the RootFS.Layers in docker inspect are digests of the uncompressed layer tars.

You can look: docker save -o x.tar IMAGE, tar -xf, then jq on manifest.json and the config blob. Pulls are per layer, so a server only downloads the layers it does not have.

Also asked: Why should you deploy by digest rather than by tag? · Why do the layer IDs in docker pull output not match docker inspect? · Why does pulling a new version of an image download only some layers?

Learn it: 10.53 Image anatomy: manifest, config, layers, digests

Describe a good image tagging strategy. Junior

Tags that never move, each traceable to one commit:

Make the registry enforce it with immutable tags, so a re-push of 2.14.1 is rejected. Record the digest docker push prints, and have production reference the digest (or an immutable tag). Then a rollback is "deploy the previous tag", and it still means what it meant yesterday.

The mechanics: docker login registry.lab --password-stdin (not -p), docker tag orders:slim registry.lab/team/orders:2.14.1, docker push.

Also asked: How do you push an image to a private registry? · What does "exec format error" mean when a container starts? · What is a multi-arch image, and how do you build one?

Learn it: 10.54 Registries, tagging strategy and multi-arch images

Practise these answers with flashcards and labs Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.