Two levels
The problem. Most public images run as root, can write anywhere in their own filesystem and keep a dozen kernel privileges they never use. If an attacker gets code running inside, all of that is theirs. A few lines in the pod spec take it away.
What you need to know already: users, uids and setuid (4.3, 4.9), systemd hardening like NoNewPrivileges= (2.26), container capabilities and read-only containers in Docker (11.29), non-root images (10.40), emptyDir (16.37), /proc/PID/status (3.14).
The securityContext is the part of a pod spec that says which user a container runs as and which privileges it keeps. It exists at two levels:
spec:
securityContext: # POD level: applies to all containers (and volumes)
runAsNonRoot: true
runAsUser: 10001
runAsGroup: 10001
fsGroup: 10001 # group owner of mounted volumes (emptyDir, PVCs)
seccompProfile:
type: RuntimeDefault
containers:
- name: app
securityContext: # CONTAINER level: overrides the pod level
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop: ["ALL"]
Some fields exist only at one level: fsGroup, supplementalGroups and sysctls are pod-level; capabilities, privileged, readOnlyRootFilesystem and allowPrivilegeEscalation are container-level. runAsUser, runAsGroup, runAsNonRoot and seccompProfile exist at both (container wins).
That block is the "restricted" baseline most platform teams require (lesson 17.39 makes the cluster enforce it). Each line closes a real hole. Short glossary of the lines: fsGroup = the group that owns mounted volumes; seccomp = a kernel filter on which system calls the process may make; capabilities = the separate pieces of root's power (11.29).
runAsNonRoot and runAsUser
By default a container runs as whatever USER its image says - and for most public images (nginx, busybox, many app images) that is root, uid 0. Root inside a container is still root on the node's kernel; a container escape from uid 0 is a node takeover.
runAsUser: 10001- run as that uid, whatever the image says.runAsNonRoot: true- the kubelet refuses to start the container unless it can prove the uid is not 0. The pod then showsCreateContainerConfigError(the kubelet could not build the container from its config):
# an illustration: pods with the securityContext above (the securityContext mission)
kubectl get pods
NAME READY STATUS RESTARTS AGE
web 0/1 CreateContainerConfigError 0 12s
kubectl describe pod web | grep -A3 'State:'
State: Waiting
Reason: CreateContainerConfigError
Message: container has runAsNonRoot and image will run as root (pod: "web_shop(2f42b38a-...)", container: web)
Other variants: container's runAsUser breaks non-root policy (you set runAsNonRoot and runAsUser: 0 together), and image has non-numeric user (nginx), cannot verify user is non-root - the image's USER is a name, and the kubelet cannot resolve names, so set a numeric runAsUser.
Inside, check who you are:
# an illustration: pods with the securityContext above (the securityContext mission)
kubectl exec box -- id
uid=1000 gid=3000 groups=2000,3000
kubectl exec box -- whoami
whoami: unknown uid 1000
command terminated with exit code 1
(No passwd entry for 1000 - harmless, but some apps and shells complain about it. fsGroup: 2000 shows up as a supplementary group, and files the kubelet writes into volumes are group-owned by it.)
Running an image that expects root as non-root usually breaks on the first write. nginx is the textbook case:
# an illustration: pods with the securityContext above (the securityContext mission)
kubectl logs web
/docker-entrypoint.sh: /docker-entrypoint.d/ is not empty, will attempt to perform configuration
...
nginx: [warn] the "user" directive makes sense only if the master process runs with super-user privileges, ignored in /etc/nginx/nginx.conf:2
nginx: [emerg] mkdir() "/var/cache/nginx/client_temp" failed (13: Permission denied)
The fix is an image built for it: nginxinc/nginx-unprivileged runs as uid 101, listens on 8080 (ports below 1024 need root or a special capability), and keeps its temp files and pid in /tmp.
readOnlyRootFilesystem
The container's own filesystem (the image layers + a writable overlay) becomes read-only. An attacker who gets code execution cannot drop a binary, edit a config or plant a cron job. Everything the app legitimately writes must go to a volume:
securityContext:
readOnlyRootFilesystem: true
volumeMounts:
- {name: tmp, mountPath: /tmp}
volumes:
- name: tmp
emptyDir: {}
Without the volume, the failure is immediate and specific:
nginx: [emerg] mkdir() "/var/cache/nginx/client_temp" failed (30: Read-only file system)
Caused by: org.springframework.boot.web.server.WebServerException: Unable to create tempDir. java.io.tmpdir is set to /tmp
(The second one is a Java/Spring Boot app failing the same way - Java apps write temp files to /tmp too.)
(13: Permission denied) = the path is writable but not by your uid. (30: Read-only file system) = nothing may write there. Two different fixes. From a shell in the container:
# an illustration: pods with the securityContext above (the securityContext mission)
kubectl exec web -- touch /etc/x
touch: /etc/x: Read-only file system
command terminated with exit code 1
kubectl exec web -- touch /tmp/x # the emptyDir: fine
Capabilities
Root's powers are split into ~40 capabilities (CAP_NET_ADMIN, CAP_SYS_ADMIN, CAP_CHOWN...). containerd gives containers a default set of 14 (CHOWN, DAC_OVERRIDE, FOWNER, FSETID, KILL, SETGID, SETUID, SETPCAP, NET_BIND_SERVICE, NET_RAW, SYS_CHROOT, MKNOD, AUDIT_WRITE, SETFCAP). Most apps need none of them:
securityContext:
capabilities:
drop: ["ALL"]
add: ["NET_BIND_SERVICE"] # only if it must bind a port < 1024 as non-root
The kernel shows the effective set of PID 1 as a bitmask:
# an illustration: pods with the securityContext above (the securityContext mission)
kubectl exec web -- grep Cap /proc/1/status
CapInh: 0000000000000000
CapPrm: 00000000a80425fb <- the default 14, running as root
CapEff: 00000000a80425fb
CapBnd: 00000000a80425fb
CapAmb: 0000000000000000
After drop: ["ALL"] every line is 0000000000000000. (A non-root process has an empty effective set anyway - CapEff 0 - but the bounding set, CapBnd, still limits what a setuid binary could regain. Dropping ALL empties that too.)
Dropping capabilities from a root nginx breaks it in yet another way - it cannot chown its temp dirs any more:
nginx: [emerg] chown("/var/cache/nginx/client_temp", 101) failed (1: Operation not permitted)
privileged: true is the opposite extreme: every capability, all devices, no isolation to speak of. Only node agents (CNI network plugins, 16.35, and storage drivers) have a reason.
allowPrivilegeEscalation and seccomp
allowPrivilegeEscalation: falsesets the kernel's no_new_privs flag: no setuid/setgid binary (sudo, su, ping on old images) can raise privileges again. It shows asNoNewPrivs: 1in/proc/1/status. (For aprivilegedcontainer or one with CAP_SYS_ADMIN, escalation is always possible whatever you set - one more reason restricted forbids both.)seccompProfile: {type: RuntimeDefault}filters system calls through the runtime's default profile (blocks ~60 rarely-needed, often-exploited syscalls likekeyctl,unshare,mount). Kubernetes does not apply it by default - pods areUnconfined(Seccomp: 0in /proc/1/status) unless you ask, or the kubelet runs with--seccomp-default. With it:Seccomp: 2(filter mode).
A hardened pod, verified
# an illustration: pods with the securityContext above (the securityContext mission)
kubectl exec web -- grep -E 'Uid|CapEff|CapBnd|NoNewPrivs|Seccomp:' /proc/1/status
Uid: 101 101 101 101
CapEff: 0000000000000000
CapBnd: 0000000000000000
NoNewPrivs: 1
Seccomp: 2
Non-root, no capabilities, cannot escalate, syscall-filtered, and (if you try to write) a read-only root. That is what the restricted Pod Security Standard requires - next lesson, where the cluster enforces it for you.
What you can now do
- Write the standard hardened securityContext and know what each line blocks.
- Read the failure each one causes (root image refused, Permission denied vs Read-only file system).
- Verify a running container from
/proc/1/status.