Why this matters
Two things turn a small problem into a big one. A container running as root: if someone breaks out of it, they are root on the host. And a password baked into an image: everyone who can download the image can read it, forever, even if a later line "deleted" it. Both are fixed in the Dockerfile.
What you need to know already: uid 0 in a container is uid 0 on the host ("What a container actually is"), whiteouts and root-owned COPY (the history lesson), ENV is not secret (Ch 2, "Prove an env var is not a secret"), tar (Ch 4), users and useradd (Ch 4).
Run as a non-root user
By default a container runs as root - and without user namespaces, that is uid 0 on the host kernel. What stops it is a set of kernel restrictions Docker turns on: dropped capabilities (root's powers split into pieces, most of them removed), seccomp (a filter that blocks dangerous system calls) and AppArmor (a policy of what files a program may touch). Useful, but a breakout from a root container is a root breakout. Run as someone else:
FROM eclipse-temurin:21-jre
RUN groupadd -r app && useradd -r -g app -u 10001 app
WORKDIR /app
COPY --chown=app:app --from=build /app/target/app.jar app.jar
USER app
ENTRYPOINT ["java","-jar","app.jar"]
# orders:secure = the image you harden in the mission below
docker run --rm orders:secure id
uid=10001(app) gid=10001(app) groups=10001(app)
groupadd -r appcreates a system group;useradd -r -g app -u 10001 appcreates a system user (-r: no home, no password aging) in that group (-g) with uid 10001 (-u). On alpine the busybox spelling isaddgroup -S app && adduser -S -G app app.USER 10001(a number) also works without an/etc/passwdentry - but then tools printI have no name!andwhoamifails. Many teams useUSER 10001:10001with the user also created: a number can be checked by tools without looking inside the image.- Anything the app must write (a cache dir, uploads, a PID file) must be owned by that user:
COPY --chown, orRUN mkdir /app/data && chown app:app /app/datafor an empty directory. Remember thatCOPYfiles are root-owned otherwise, whateverUSERsays. USERapplies to every laterRUN, so put it after theapt-gets:
> [4/5] RUN apt-get update:
0.412 E: List directory /var/lib/apt/lists/partial is missing. - Acquire (13: Permission denied)
Secrets: ARG and ENV are both visible
A secret is anything that grants access: a password, a token, a private key. Here a token for a private npm registry, NPM_TOKEN, is passed as an ARG, and a database password as an ENV:
FROM node:22-slim
ARG NPM_TOKEN
ENV DB_PASSWORD=Sup3r-s3cret-prod
RUN echo "//registry.npmjs.org/:_authToken=${NPM_TOKEN}" > ~/.npmrc && npm ci && rm ~/.npmrc
(~/.npmrc is npm's settings file, where it reads the token.) BuildKit warns twice (SecretsUsedInArgOrEnv) and it is right both times:
# grep -E 'NPM_TOKEN|DB_PASSWORD'|app:1 = the leaky image from the mission below
docker history --no-trunc app:1 | grep -E 'NPM_TOKEN|DB_PASSWORD'
<missing> RUN |1 NPM_TOKEN=npm_3kf9Qz1xT7 /bin/sh -c echo "//registry.npmjs.org/:_authToken=${NPM_TOKEN}" > ~/.npmrc && npm ci && rm ~/.npmrc # buildkit
<missing> ENV DB_PASSWORD=Sup3r-s3cret-prod
docker inspect -f '{{json .Config.Env}}' app:1
["PATH=/usr/local/sbin:...","DB_PASSWORD=Sup3r-s3cret-prod"]
- ENV is part of the image config. Every container, every
inspect, every registry that stores the image has it - the same lesson as systemd'sEnvironment=in Ch 2. - ARG is not in the final environment, but every RUN that has it in scope records it in the history as
|1 NPM_TOKEN=....
Deleting a copied secret does not remove it
COPY id_rsa /root/.ssh/id_rsa
RUN git clone [email protected]:ing/private-lib.git && rm /root/.ssh/id_rsa
(id_rsa is an SSH private key, like the ed25519 key you made in Ch 1.) The final filesystem has no key. The COPY layer still does - the later layer only adds a whiteout. You can prove it without any special tools:
# same app:1; blob names shortened
docker save -o app.tar app:1
mkdir app-img && tar -xf app.tar -C app-img
ls app-img
blobs index.json manifest.json oci-layout repositories
cat app-img/manifest.json
[{"Config":"blobs/sha256/5d1f...","RepoTags":["app:1"],"Layers":["blobs/sha256/0879...","blobs/sha256/64c3...","blobs/sha256/9a1b...","blobs/sha256/c2d4..."]}]
tar -tvf app-img/blobs/sha256/c2d4...
-rw-r--r-- 0/0 0 2026-09-23 10:00 root/.ssh/.wh.id_rsa
tar -tvf app-img/blobs/sha256/9a1b...
-rw------- 0/0 2602 2026-09-23 10:00 root/.ssh/id_rsa
tar -xOf app-img/blobs/sha256/9a1b... root/.ssh/id_rsa
-----BEGIN OPENSSH PRIVATE KEY-----
The commands: docker save -o app.tar app:1 writes the whole image to one tar file. tar -xf FILE -C DIR extracts into DIR; tar -tvf lists an archive's contents in detail; tar -xOf ARCHIVE PATH extracts one file to stdout (-O). Each blob (a file named by its sha256) is one layer, or the config; manifest.json lists them in order.
.wh.id_rsa is the whiteout: the marker a later layer uses to hide a file. The file itself is one layer down, intact. The config blob (the Config entry in manifest.json) is JSON with the full history - ARGs included. Anyone who can docker pull your image can do this. (dive is a popular terminal tool that shows the same thing layer by layer.)
The fix: secret mounts
BuildKit can mount a secret into one RUN, as a file that never lands in a layer or in the history:
# syntax=docker/dockerfile:1
FROM node:22-slim
WORKDIR /app
COPY package.json package-lock.json ./
RUN --mount=type=secret,id=npmrc,target=/root/.npmrc npm ci
docker build --secret id=npmrc,src=$HOME/.npmrc -t app:1 .
--secret id=npmrc,src=FILE hands the file to the build under the name npmrc; --mount=type=secret,id=npmrc,target=PATH makes it appear at PATH during that one RUN only.
# grep npm|after the rebuild with a secret mount
docker history --no-trunc app:1 | grep npm
<missing> RUN /bin/sh -c npm ci # buildkit
No token in the history, no .npmrc in any layer. Without target= the secret appears as /run/secrets/<id>; env=NAME exposes it as an environment variable for that RUN only. For SSH keys there is --mount=type=ssh plus docker build --ssh default, which lends your SSH agent to the build instead of copying a key.
Runtime secrets are a different problem
A database password the running app needs does not belong in the image at all. Hand it over when the container starts: -e NAME=value / --env-file FILE (visible in docker inspect), or better, a mounted read-only file (-v /etc/orders/db.pass:/run/secrets/db.pass:ro, :ro = read-only) - the same idea as systemd's LoadCredential= in Ch 2. The rule for images is simple: an image should be safe to make public. If it is not, something is baked in that should not be.
What you can now do
- Add a non-root user and own the writable paths correctly.
- Prove a secret leaked with
docker historyanddocker save+tar. - Pass a build secret with
--mount=type=secretinstead.