Why your pod is not running as the user you asked for
You deploy an image whose Dockerfile says USER 101. On OpenShift, id inside the pod says uid=1000650000. Nobody changed your YAML - OpenShift changed the pod on its way in. Until you know what it changes and why, every "Permission denied" here is a mystery.
What you need to know already: admission, the step where the apiserver may change or reject an object (15.5); a pod's securityContext: runAsUser, runAsNonRoot, fsGroup, capabilities, allowPrivilegeEscalation, seccompProfile (17.37); Pod Security Admission and its restricted level (17.39); capabilities (11.29); ClusterRole / RoleBinding and auth can-i (17.30, 17.35); projects and their UID range annotation (30.4).
The words you need first
- SCC (SecurityContextConstraints) - an OpenShift object that lists what a pod may ask for (root? host network? which volumes?) and what it gets when it does not ask.
- Strategy - an SCC's rule for one field, e.g. "the UID must come from the project's range" (
MustRunAsRange) or "any UID is fine" (RunAsAny). - Mutate / validate - admission can mutate an object (fill in or change fields) and then validate it (accept or reject). PSA only validates; SCC does both.
An admission plugin with opinions
A SecurityContextConstraints object (SCC, API group security.openshift.io, cluster-scoped) describes what a pod is allowed to ask for and what it gets when it does not ask. Every pod OpenShift admits passes through the SCC admission plugin, which:
- collects the SCCs the pod may use (next lesson: who may use what),
- tries them in order,
- for each one fills in the fields the pod left empty (UID, fsGroup, SELinux level, seccomp, dropped capabilities...) and then validates the result,
- admits the pod with the first SCC that validates, and writes its name into the pod's
openshift.io/sccannotation - or rejects the pod with every SCC's reason.
Step 3 is the part Kubernetes has no equivalent for. Pod Security Admission only judges a pod; SCC admission also rewrites it. That is why the same YAML behaves differently here.
The SCCs a 4.21 cluster ships
oc get scc lists them. The table is wide; the next paragraph explains the columns you care about, so do not try to read every cell yet:
# cluster-scoped: as kubeadmin
oc get scc
NAME PRIV CAPS SELINUX RUNASUSER FSGROUP SUPGROUP PRIORITY READONLYROOTFS VOLUMES
anyuid false <no value> MustRunAs RunAsAny RunAsAny RunAsAny 10 false ["configMap","csi","downwardAPI","emptyDir","ephemeral","persistentVolumeClaim","projected","secret"]
hostaccess false <no value> MustRunAs MustRunAsRange MustRunAs RunAsAny <no value> false ["configMap","csi","downwardAPI","emptyDir","ephemeral","hostPath","persistentVolumeClaim","projected","secret"]
hostmount-anyuid false <no value> MustRunAs RunAsAny RunAsAny RunAsAny <no value> false ["configMap","csi","downwardAPI","emptyDir","ephemeral","hostPath","nfs","persistentVolumeClaim","projected","secret"]
hostnetwork false <no value> MustRunAs MustRunAsRange MustRunAs MustRunAs <no value> false [...]
hostnetwork-v2 false ["NET_BIND_SERVICE"] MustRunAs MustRunAsRange MustRunAs MustRunAs <no value> false [...]
nested-container false <no value> MustRunAs MustRunAsRange MustRunAs RunAsAny <no value> false [...]
node-exporter true <no value> RunAsAny RunAsAny RunAsAny RunAsAny <no value> false ["*"]
nonroot false <no value> MustRunAs MustRunAsNonRoot RunAsAny RunAsAny <no value> false [...]
nonroot-v2 false ["NET_BIND_SERVICE"] MustRunAs MustRunAsNonRoot RunAsAny RunAsAny <no value> false [...]
privileged true ["*"] RunAsAny RunAsAny RunAsAny RunAsAny <no value> false ["*"]
restricted false <no value> MustRunAs MustRunAsRange MustRunAs RunAsAny <no value> false [...]
restricted-v2 false ["NET_BIND_SERVICE"] MustRunAs MustRunAsRange MustRunAs RunAsAny <no value> false [...]
restricted-v3 false ["NET_BIND_SERVICE"] MustRunAs MustRunAsRange MustRunAs RunAsAny <no value> false [...]
(Listing SCCs needs cluster rights; as developer you get Error from server (Forbidden): securitycontextconstraints.security.openshift.io is forbidden.)
Read the columns as "strategies":
- RUNASUSER -
MustRunAsRange= a UID from the project's range, defaulting to its first value;MustRunAsNonRoot= any UID except 0, and you must supply it (or the image's numeric USER must be non-zero);RunAsAny= whatever the pod or image says, root included. - SELINUX -
MustRunAs= the project's SELinux MCS level (openshift.io/sa.scc.mcs, 30.4). - FSGROUP / SUPGROUP -
MustRunAs= from the project's supplemental-groups range. - CAPS - capabilities a pod may add. The
-v2SCCs also drop ALL and allow adding back onlyNET_BIND_SERVICE. - PRIV -
allowPrivilegedContainer: may a container runprivileged: true(nearly all of root's powers on the node)? VOLUMES - allowed volume types (16.37); nohostPath(a folder of the node itself) outside the host* SCCs. - PRIORITY -
anyuidhas 10, so for anyone allowed to use it, it is tried first.
What each one is for, in one line:
| SCC | for |
|---|---|
restricted-v2 | every authenticated user and ServiceAccount; the default for normal pods |
restricted-v3 | 4.20+: like restricted-v2 but the pod must run in a user namespace (hostUsers: false: UID 0 inside the container maps to an unprivileged UID on the node) |
nonroot-v2 | a fixed non-root UID of your choosing (runAsUser: 1001) instead of the project's |
anyuid | any UID, root included. Priority 10. Granted to cluster admins |
hostnetwork-v2 | hostNetwork / hostPort, still non-root |
hostmount-anyuid, hostaccess | host paths / host namespaces - node agents |
privileged | everything. The platform's own daemons; nobody else |
restricted, nonroot, hostnetwork | the pre-4.11 versions; on fresh 4.11+ installs restricted is not granted to users any more |
nested-container | running podman/buildah (Docker-like tools that need no daemon) inside a pod |
restricted-v3: whether ordinary users are granted it depends on the version and how the cluster was upgraded - check yours with oc describe clusterrolebinding system:openshift:scc:restricted-v3 (or oc adm policy who-can use scc restricted-v3). Either way, a pod that does not set hostUsers: false fails its check and lands on restricted-v2, which is what you will see on almost every pod.
restricted-v2 in detail
oc describe scc NAME prints one SCC as a readable list. Most lines are "no" to something (no privileged, no host network, no host ports); the four Strategy lines at the bottom are the ones that change your pod:
# as kubeadmin
oc describe scc restricted-v2
Name: restricted-v2
Priority: <none>
Access:
Users: <none>
Groups: <none>
Settings:
Allow Privileged: false
Allow Privilege Escalation: false
Default Add Capabilities: <none>
Required Drop Capabilities: ALL
Allowed Capabilities: NET_BIND_SERVICE
Allowed Seccomp Profiles: runtime/default
Allowed Volume Types: configMap,csi,downwardAPI,emptyDir,ephemeral,persistentVolumeClaim,projected,secret
Allowed Flexvolumes: <all>
Allowed Unsafe Sysctls: <none>
Forbidden Sysctls: <none>
Allow Host Network: false
Allow Host Ports: false
Allow Host PID: false
Allow Host IPC: false
Read Only Root Filesystem: false
Run As User Strategy: MustRunAsRange
UID: <none>
UID Range Min: <none>
UID Range Max: <none>
SELinux Context Strategy: MustRunAs
User: <none>
Role: <none>
Type: <none>
Level: <none>
FSGroup Strategy: MustRunAs
Ranges: <none>
Supplemental Groups Strategy: RunAsAny
Ranges: <none>
"Users: <none>, Groups: <none>" does not mean nobody can use it. oc describe scc only shows the SCC's own users/groups fields. Access is granted through RBAC: the ClusterRole system:openshift:scc:restricted-v2 (verb use on that SCC) is bound to the group system:authenticated. The "UID Range Min/Max: <none>" means "take it from the namespace annotation".
Before and after
You apply a Deployment with no securityContext at all:
containers:
- name: web
image: docker.io/nginxinc/nginx-unprivileged:1.27
Creating it already prints a warning - Pod Security Admission, not SCC:
Warning: would violate PodSecurity "restricted:latest": allowPrivilegeEscalation != false (container "web" must set securityContext.allowPrivilegeEscalation=false), unrestricted capabilities (container "web" must set securityContext.capabilities.drop=["ALL"]), runAsNonRoot != true (pod or container "web" must set securityContext.runAsNonRoot=true), seccompProfile (pod or container "web" must set securityContext.seccompProfile.type to "RuntimeDefault" or "Localhost")
deployment.apps/web created
The project's PSA warn label is restricted (30.4), and PSA looks at the Deployment's template, which really does lack those fields. The warning is harmless noise: the pods themselves will be fixed up by SCC admission before PSA sees them. Teams that want a quiet oc apply put those four fields in their templates - which is also exactly what makes the manifest portable to a strict PSA cluster.
The pod that actually runs (grep -B2 -A12 shows 2 lines before and 12 after each match, 7.1):
# web = the deployment of the next mission
oc get pod -l app=web -o yaml | grep -B2 -A12 'securityContext:'
securityContext:
allowPrivilegeEscalation: false
capabilities:
drop:
- ALL
runAsNonRoot: true
runAsUser: 1000650000
...
securityContext:
fsGroup: 1000650000
seLinuxOptions:
level: s0:c27,c14
seccompProfile:
type: RuntimeDefault
Everything there was written by admission: the first UID and group of the project range, the MCS level, drop ALL, no escalation, the RuntimeDefault seccomp profile. And admission leaves its signature in an annotation. The jsonpath (15.38) needs \. because the annotation name itself contains dots; oc rsh deploy/web id runs id in one of the Deployment's pods:
oc get pod -l app=web -o jsonpath='{.items[0].metadata.annotations.openshift\.io/scc}{"\n"}'
restricted-v2
oc rsh deploy/web id
uid=1000650000(1000650000) gid=0(root) groups=0(root),1000650000
(the simulator prints uid=1000650000 gid=0(root)...: there is no passwd entry for that UID, so real id also shows just the number unless the image's entrypoint adds one.)
Three facts to take away from that one line:
- The UID is not the image's USER. Whatever
USERthe Dockerfile set - root, 101, 1001 - restricted-v2 replaces it with the project's UID. Different project, different UID; you cannot know it at build time. - The primary group is 0 (root). Not root user - root group. This is the hook that makes arbitrary UIDs workable: an image that makes its writable directories owned by group 0 and group-writable works for any UID.
- fsGroup is also added (
groups=...,1000650000), so volumes that support ownership management (emptyDir, most block PVCs) are made writable for the pod.
Pod Security Admission is still there
OpenShift runs PSA too, independently. Globally it enforces privileged (nothing blocked) and warns/audits restricted. On your projects a controller syncs the warn and audit labels to match the most privileged SCC your ServiceAccounts can use (restricted-v2 -> restricted, anyuid -> baseline, privileged -> privileged), so that you are not warned about things SCC will let through anyway. Setting a PSA label by hand switches the sync off for that label; security.openshift.io/scc.podSecurityLabelSync=false switches it off for the namespace. Enforcement on OpenShift is SCC's job.
The questions to ask of a pod on OpenShift
- Which SCC admitted it? (
openshift.io/scc) - Which UID does it actually run as? (
oc rsh ... id, or therunAsUseradmission wrote) - Does the image write outside a mounted volume, need root, chown files, or bind a port below 1024? If yes, restricted-v2 will break it - lesson 30.9.
What you can now do
- Explain what SCC admission does that Pod Security Admission does not: it rewrites the pod, then checks it.
- Read
oc get scc/oc describe sccand say what restricted-v2 gives a pod. - Find which SCC admitted a pod and which UID it really runs as.