OnCallReady

Lesson 17.37 · Kubernetes: Scheduling, Health & Security · 20 min read

securityContext: the fields worth setting every time

In plain words

Imagine giving a babysitter the keys to your house. You could give them the master key that opens everything, including the safe and the garage with the car. Or you could give them only the front door key, lock the cupboards they don't need, tell them they may not make copies of keys, and ask them not to rearrange the furniture. If something goes wrong, the damage is small.

A securityContext is that set of rules for a container. runAsNonRoot and runAsUser stop it running as root; capabilities: drop: ["ALL"] takes away root's special powers; allowPrivilegeEscalation: false stops setuid tricks; readOnlyRootFilesystem stops it writing to its own files; seccompProfile: RuntimeDefault blocks dangerous system calls. Each line closes a real hole.

Two levels

The problem. Most public images run as root, can write anywhere in their own filesystem and keep a dozen kernel privileges they never use. If an attacker gets code running inside, all of that is theirs. A few lines in the pod spec take it away.

What you need to know already: users, uids and setuid (4.3, 4.9), systemd hardening like NoNewPrivileges= (2.26), container capabilities and read-only containers in Docker (11.29), non-root images (10.40), emptyDir (16.37), /proc/PID/status (3.14).

The securityContext is the part of a pod spec that says which user a container runs as and which privileges it keeps. It exists at two levels:

spec:
  securityContext:                  # POD level: applies to all containers (and volumes)
    runAsNonRoot: true
    runAsUser: 10001
    runAsGroup: 10001
    fsGroup: 10001                  # group owner of mounted volumes (emptyDir, PVCs)
    seccompProfile:
      type: RuntimeDefault
  containers:
  - name: app
    securityContext:                # CONTAINER level: overrides the pod level
      allowPrivilegeEscalation: false
      readOnlyRootFilesystem: true
      capabilities:
        drop: ["ALL"]

Some fields exist only at one level: fsGroup, supplementalGroups and sysctls are pod-level; capabilities, privileged, readOnlyRootFilesystem and allowPrivilegeEscalation are container-level. runAsUser, runAsGroup, runAsNonRoot and seccompProfile exist at both (container wins).

That block is the "restricted" baseline most platform teams require (lesson 17.39 makes the cluster enforce it). Each line closes a real hole. Short glossary of the lines: fsGroup = the group that owns mounted volumes; seccomp = a kernel filter on which system calls the process may make; capabilities = the separate pieces of root's power (11.29).

runAsNonRoot and runAsUser

By default a container runs as whatever USER its image says - and for most public images (nginx, busybox, many app images) that is root, uid 0. Root inside a container is still root on the node's kernel; a container escape from uid 0 is a node takeover.

# an illustration: pods with the securityContext above (the securityContext mission)
kubectl get pods
NAME   READY   STATUS                       RESTARTS   AGE
web    0/1     CreateContainerConfigError   0          12s
kubectl describe pod web | grep -A3 'State:'
    State:          Waiting
      Reason:       CreateContainerConfigError
      Message:      container has runAsNonRoot and image will run as root (pod: "web_shop(2f42b38a-...)", container: web)

Other variants: container's runAsUser breaks non-root policy (you set runAsNonRoot and runAsUser: 0 together), and image has non-numeric user (nginx), cannot verify user is non-root - the image's USER is a name, and the kubelet cannot resolve names, so set a numeric runAsUser.

Inside, check who you are:

# an illustration: pods with the securityContext above (the securityContext mission)
kubectl exec box -- id
uid=1000 gid=3000 groups=2000,3000
kubectl exec box -- whoami
whoami: unknown uid 1000
command terminated with exit code 1

(No passwd entry for 1000 - harmless, but some apps and shells complain about it. fsGroup: 2000 shows up as a supplementary group, and files the kubelet writes into volumes are group-owned by it.)

Running an image that expects root as non-root usually breaks on the first write. nginx is the textbook case:

# an illustration: pods with the securityContext above (the securityContext mission)
kubectl logs web
/docker-entrypoint.sh: /docker-entrypoint.d/ is not empty, will attempt to perform configuration
...
nginx: [warn] the "user" directive makes sense only if the master process runs with super-user privileges, ignored in /etc/nginx/nginx.conf:2
nginx: [emerg] mkdir() "/var/cache/nginx/client_temp" failed (13: Permission denied)

The fix is an image built for it: nginxinc/nginx-unprivileged runs as uid 101, listens on 8080 (ports below 1024 need root or a special capability), and keeps its temp files and pid in /tmp.

readOnlyRootFilesystem

The container's own filesystem (the image layers + a writable overlay) becomes read-only. An attacker who gets code execution cannot drop a binary, edit a config or plant a cron job. Everything the app legitimately writes must go to a volume:

    securityContext:
      readOnlyRootFilesystem: true
    volumeMounts:
    - {name: tmp, mountPath: /tmp}
  volumes:
  - name: tmp
    emptyDir: {}

Without the volume, the failure is immediate and specific:

nginx: [emerg] mkdir() "/var/cache/nginx/client_temp" failed (30: Read-only file system)
Caused by: org.springframework.boot.web.server.WebServerException: Unable to create tempDir. java.io.tmpdir is set to /tmp

(The second one is a Java/Spring Boot app failing the same way - Java apps write temp files to /tmp too.)

(13: Permission denied) = the path is writable but not by your uid. (30: Read-only file system) = nothing may write there. Two different fixes. From a shell in the container:

# an illustration: pods with the securityContext above (the securityContext mission)
kubectl exec web -- touch /etc/x
touch: /etc/x: Read-only file system
command terminated with exit code 1
kubectl exec web -- touch /tmp/x          # the emptyDir: fine

Capabilities

Root's powers are split into ~40 capabilities (CAP_NET_ADMIN, CAP_SYS_ADMIN, CAP_CHOWN...). containerd gives containers a default set of 14 (CHOWN, DAC_OVERRIDE, FOWNER, FSETID, KILL, SETGID, SETUID, SETPCAP, NET_BIND_SERVICE, NET_RAW, SYS_CHROOT, MKNOD, AUDIT_WRITE, SETFCAP). Most apps need none of them:

    securityContext:
      capabilities:
        drop: ["ALL"]
        add: ["NET_BIND_SERVICE"]      # only if it must bind a port < 1024 as non-root

The kernel shows the effective set of PID 1 as a bitmask:

# an illustration: pods with the securityContext above (the securityContext mission)
kubectl exec web -- grep Cap /proc/1/status
CapInh:	0000000000000000
CapPrm:	00000000a80425fb           <- the default 14, running as root
CapEff:	00000000a80425fb
CapBnd:	00000000a80425fb
CapAmb:	0000000000000000

After drop: ["ALL"] every line is 0000000000000000. (A non-root process has an empty effective set anyway - CapEff 0 - but the bounding set, CapBnd, still limits what a setuid binary could regain. Dropping ALL empties that too.)

Dropping capabilities from a root nginx breaks it in yet another way - it cannot chown its temp dirs any more:

nginx: [emerg] chown("/var/cache/nginx/client_temp", 101) failed (1: Operation not permitted)

privileged: true is the opposite extreme: every capability, all devices, no isolation to speak of. Only node agents (CNI network plugins, 16.35, and storage drivers) have a reason.

allowPrivilegeEscalation and seccomp

A hardened pod, verified

# an illustration: pods with the securityContext above (the securityContext mission)
kubectl exec web -- grep -E 'Uid|CapEff|CapBnd|NoNewPrivs|Seccomp:' /proc/1/status
Uid:	101	101	101	101
CapEff:	0000000000000000
CapBnd:	0000000000000000
NoNewPrivs:	1
Seccomp:	2

Non-root, no capabilities, cannot escalate, syscall-filtered, and (if you try to write) a read-only root. That is what the restricted Pod Security Standard requires - next lesson, where the cluster enforces it for you.

What you can now do

Why it helps

Every workload on a bank's platform goes through a security review, and this block is what the reviewer checks. Root in a container is root on the node's kernel, so a container escape from uid 0 is a node takeover; the restricted Pod Security level requires most of these fields, so teams will come to you when their pod is rejected or won't start.

You'll also debug the fallout: CreateContainerConfigError with "image will run as root", nginx failing with Permission denied (wrong uid) versus Read-only file system (needs an emptyDir), Spring Boot unable to create its temp dir. Knowing which error means which fix, and reading /proc/1/status to verify (CapEff, NoNewPrivs, Seccomp), makes you the person who unblocks teams instead of just rejecting them.

FAQ

Isn't root inside a container harmless because it's isolated?

No. Containers share the node's kernel, and uid 0 inside is uid 0 to that kernel, only limited by namespaces, capabilities and seccomp. A kernel bug or misconfiguration (a hostPath, a privileged flag, a mounted runtime socket) turns root in the container into root on the node. Running as non-root removes a whole class of escalation paths.

Why does my pod fail with "image will run as root"?

You set runAsNonRoot: true and the image's USER is root (or unset). The kubelet refuses to start the container: CreateContainerConfigError. Set a numeric runAsUser, or use an image built to run unprivileged. If the image's USER is a name like nginx, the kubelet can't verify it's non-root, so a numeric runAsUser is needed there too.

What's the difference between "Permission denied" and "Read-only file system"?

Errno 13, Permission denied, means the path is writable, just not by your uid: fix ownership, use fsGroup for volumes, or use an image built for that uid. Errno 30, Read-only file system, means nothing may write there because readOnlyRootFilesystem is on: mount an emptyDir at the paths the app writes, like /tmp or a cache directory.

Does Kubernetes apply seccomp by default?

No. Pods run Unconfined unless you set seccompProfile: {type: RuntimeDefault} or the kubelet runs with --seccomp-default. RuntimeDefault blocks a set of rarely needed, often exploited syscalls like keyctl, unshare and mount. grep Seccomp /proc/1/status shows 0 when unconfined and 2 when filtering.

Why drop ALL capabilities if the container runs as non-root anyway?

A non-root process has an empty effective set, but the bounding set still defines what a setuid binary could regain. drop: ["ALL"] empties the bounding set too, and together with allowPrivilegeEscalation: false (no_new_privs) closes that path. Add back only what's needed, typically just NET_BIND_SERVICE to bind a port below 1024.

In an interview Mid

Which securityContext settings would you require for every application container, and why?

Each closes a real hole if someone gets code running in the container:

And never privileged: true for an app. Verify from /proc/1/status (Uid, CapEff, NoNewPrivs, Seccomp). The failures tell you what to fix: Permission denied (wrong uid) vs Read-only file system (needs a volume). That block is exactly the restricted Pod Security Standard.

Also asked: A container fails after you enforced readOnlyRootFilesystem and runAsNonRoot. How do you help the team? · What is a Linux capability, and why drop them all? · What does allowPrivilegeEscalation: false do?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.