OnCallReady

Lesson 10.40 · Images & Builds · 17 min read

Non-root users, and secrets that leak into images

In plain words

Imagine writing your house key's secret code on a sticky note, putting it inside a photo album, and then covering it with another photo. When you give copies of the album to friends, anyone can lift the photo and read the code. Covering is not removing.

Image layers work like that album. A key copied in one layer and deleted in the next is still in the first layer, hidden by a whiteout. ENV values sit in the config for anyone who runs docker inspect, and ARG values show up in docker history. The fix is a secret mount, RUN --mount=type=secret, which lends the secret to one step without writing it anywhere. And the process should not run as root: create a user and set USER.

Why this matters

Two things turn a small problem into a big one. A container running as root: if someone breaks out of it, they are root on the host. And a password baked into an image: everyone who can download the image can read it, forever, even if a later line "deleted" it. Both are fixed in the Dockerfile.

What you need to know already: uid 0 in a container is uid 0 on the host ("What a container actually is"), whiteouts and root-owned COPY (the history lesson), ENV is not secret (Ch 2, "Prove an env var is not a secret"), tar (Ch 4), users and useradd (Ch 4).

Run as a non-root user

By default a container runs as root - and without user namespaces, that is uid 0 on the host kernel. What stops it is a set of kernel restrictions Docker turns on: dropped capabilities (root's powers split into pieces, most of them removed), seccomp (a filter that blocks dangerous system calls) and AppArmor (a policy of what files a program may touch). Useful, but a breakout from a root container is a root breakout. Run as someone else:

FROM eclipse-temurin:21-jre
RUN groupadd -r app && useradd -r -g app -u 10001 app
WORKDIR /app
COPY --chown=app:app --from=build /app/target/app.jar app.jar
USER app
ENTRYPOINT ["java","-jar","app.jar"]
# orders:secure = the image you harden in the mission below
docker run --rm orders:secure id
uid=10001(app) gid=10001(app) groups=10001(app)
 > [4/5] RUN apt-get update:
0.412 E: List directory /var/lib/apt/lists/partial is missing. - Acquire (13: Permission denied)

Secrets: ARG and ENV are both visible

A secret is anything that grants access: a password, a token, a private key. Here a token for a private npm registry, NPM_TOKEN, is passed as an ARG, and a database password as an ENV:

FROM node:22-slim
ARG NPM_TOKEN
ENV DB_PASSWORD=Sup3r-s3cret-prod
RUN echo "//registry.npmjs.org/:_authToken=${NPM_TOKEN}" > ~/.npmrc && npm ci && rm ~/.npmrc

(~/.npmrc is npm's settings file, where it reads the token.) BuildKit warns twice (SecretsUsedInArgOrEnv) and it is right both times:

#  grep -E 'NPM_TOKEN|DB_PASSWORD'|app:1 = the leaky image from the mission below
docker history --no-trunc app:1 | grep -E 'NPM_TOKEN|DB_PASSWORD'
<missing>  RUN |1 NPM_TOKEN=npm_3kf9Qz1xT7 /bin/sh -c echo "//registry.npmjs.org/:_authToken=${NPM_TOKEN}" > ~/.npmrc && npm ci && rm ~/.npmrc # buildkit
<missing>  ENV DB_PASSWORD=Sup3r-s3cret-prod
docker inspect -f '{{json .Config.Env}}' app:1
["PATH=/usr/local/sbin:...","DB_PASSWORD=Sup3r-s3cret-prod"]

Deleting a copied secret does not remove it

COPY id_rsa /root/.ssh/id_rsa
RUN git clone [email protected]:ing/private-lib.git && rm /root/.ssh/id_rsa

(id_rsa is an SSH private key, like the ed25519 key you made in Ch 1.) The final filesystem has no key. The COPY layer still does - the later layer only adds a whiteout. You can prove it without any special tools:

# same app:1; blob names shortened
docker save -o app.tar app:1
mkdir app-img && tar -xf app.tar -C app-img
ls app-img
blobs  index.json  manifest.json  oci-layout  repositories
cat app-img/manifest.json
[{"Config":"blobs/sha256/5d1f...","RepoTags":["app:1"],"Layers":["blobs/sha256/0879...","blobs/sha256/64c3...","blobs/sha256/9a1b...","blobs/sha256/c2d4..."]}]
tar -tvf app-img/blobs/sha256/c2d4...
-rw-r--r-- 0/0            0 2026-09-23 10:00 root/.ssh/.wh.id_rsa
tar -tvf app-img/blobs/sha256/9a1b...
-rw------- 0/0         2602 2026-09-23 10:00 root/.ssh/id_rsa
tar -xOf app-img/blobs/sha256/9a1b... root/.ssh/id_rsa
-----BEGIN OPENSSH PRIVATE KEY-----

The commands: docker save -o app.tar app:1 writes the whole image to one tar file. tar -xf FILE -C DIR extracts into DIR; tar -tvf lists an archive's contents in detail; tar -xOf ARCHIVE PATH extracts one file to stdout (-O). Each blob (a file named by its sha256) is one layer, or the config; manifest.json lists them in order.

.wh.id_rsa is the whiteout: the marker a later layer uses to hide a file. The file itself is one layer down, intact. The config blob (the Config entry in manifest.json) is JSON with the full history - ARGs included. Anyone who can docker pull your image can do this. (dive is a popular terminal tool that shows the same thing layer by layer.)

The fix: secret mounts

BuildKit can mount a secret into one RUN, as a file that never lands in a layer or in the history:

# syntax=docker/dockerfile:1
FROM node:22-slim
WORKDIR /app
COPY package.json package-lock.json ./
RUN --mount=type=secret,id=npmrc,target=/root/.npmrc npm ci
docker build --secret id=npmrc,src=$HOME/.npmrc -t app:1 .

--secret id=npmrc,src=FILE hands the file to the build under the name npmrc; --mount=type=secret,id=npmrc,target=PATH makes it appear at PATH during that one RUN only.

#  grep npm|after the rebuild with a secret mount
docker history --no-trunc app:1 | grep npm
<missing>  RUN /bin/sh -c npm ci # buildkit

No token in the history, no .npmrc in any layer. Without target= the secret appears as /run/secrets/<id>; env=NAME exposes it as an environment variable for that RUN only. For SSH keys there is --mount=type=ssh plus docker build --ssh default, which lends your SSH agent to the build instead of copying a key.

Runtime secrets are a different problem

A database password the running app needs does not belong in the image at all. Hand it over when the container starts: -e NAME=value / --env-file FILE (visible in docker inspect), or better, a mounted read-only file (-v /etc/orders/db.pass:/run/secrets/db.pass:ro, :ro = read-only) - the same idea as systemd's LoadCredential= in Ch 2. The rule for images is simple: an image should be safe to make public. If it is not, something is baked in that should not be.

What you can now do

Why it helps

This is the lesson behind real incidents: a token in a public image on Docker Hub, a database password found by a scanner in docker inspect, an SSH key recovered from a 'deleted' layer. When it happens, you know how to prove it (docker history --no-trunc, docker save and tar) and that the fix is rotation, not a rebuild. In reviews you catch ARG NPM_TOKEN, ENV DB_PASSWORD and COPY id_rsa on sight, and you know --mount=type=secret is the replacement. For non-root, many platforms refuse to run images that start as root or use a non-numeric USER, so a Dockerfile you approve today decides whether a later deploy works. And 'how do you handle secrets in Docker builds?' is a very common interview question.

FAQ

If I delete a secret in a later layer, is it gone?

No. The later layer only contains a whiteout file, like .wh.id_rsa, which hides the file from the merged view. The original bytes are still in the earlier layer, which is stored in the registry and downloaded with every pull. docker save plus tar recovers it in a minute. If a secret was ever in a layer, rotate it.

Is ARG safe for secrets because it is not in the final environment?

No. ARG values are not in the container's environment, but every RUN that has the ARG in scope records it in the image history as |1 NAME=value, visible with docker history --no-trunc and stored in the config blob. BuildKit warns with SecretsUsedInArgOrEnv. Use --mount=type=secret for build secrets.

How does --mount=type=secret work?

You pass the secret to the build with docker build --secret id=npmrc,src=$HOME/.npmrc, and a RUN mounts it with --mount=type=secret,id=npmrc,target=/root/.npmrc. The file exists only while that one RUN executes. It is not written to any layer and not recorded in the history. Without target it appears at /run/secrets/<id>; env=NAME exposes it as an environment variable for that RUN only.

Why should USER be numeric?

A tool that checks whether an image runs as root, without reading /etc/passwd inside the image, can only trust a number: a name like app could map to uid 0 in that file. So platforms that enforce non-root accept only a numeric USER, and refuse USER app. Many teams create the user and then write USER 10001:10001, which gives both a name for tools and a number for checks.

Where should a runtime secret like a database password go?

Not in the image at all. Hand it over when the container starts: -e or --env-file (visible in docker inspect, so not ideal), or better a mounted read-only file such as -v /etc/orders/db.pass:/run/secrets/db.pass:ro, the same idea as systemd's LoadCredential=. The rule: an image should be safe to make public.

In an interview Junior

Why should containers not run as root, and how do you change it?

Without user namespaces, root in the container is uid 0 on the host kernel. Dropped capabilities, seccomp and AppArmor limit it, but if anything breaks out of a root container, it is root on the host.

In the Dockerfile:

RUN groupadd -r app && useradd -r -g app -u 10001 app
COPY --chown=app:app target/app.jar /app/app.jar
USER 10001

Check it: docker run --rm IMAGE id.

Also asked: How do you pass a private registry token to a build without leaking it into the image? · Why does deleting a copied secret in a later RUN not remove it from the image? · Are ARG and ENV values visible in a built image?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.