OnCallReady

Lesson 30.7 · OpenShift · 18 min read

SecurityContextConstraints I: what restricted-v2 does to your pod

In plain words

Imagine a swimming pool with a strict lifeguard. Every kid who walks in gets a wristband with a random number, has to leave the diving board alone, and cannot go into the staff room, whatever their ticket says. Kids who didn't ask for anything are simply given the standard wristband and rules. The lifeguard does not just check tickets; she changes what you are wearing at the gate.

SecurityContextConstraints are the lifeguard. The default, restricted-v2, fills in every pod's securityContext: runAsUser from the project range like 1000650000, group 0, fsGroup (the group that owns its volumes), the SELinux level (a label that keeps one project's files away from another's), drop ALL capabilities, no privilege escalation, RuntimeDefault seccomp (the runtime's default filter of blocked system calls). The pod's openshift.io/scc annotation records which SCC admitted it, and oc rsh deploy/web id shows the UID it really got.

Why your pod is not running as the user you asked for

You deploy an image whose Dockerfile says USER 101. On OpenShift, id inside the pod says uid=1000650000. Nobody changed your YAML - OpenShift changed the pod on its way in. Until you know what it changes and why, every "Permission denied" here is a mystery.

What you need to know already: admission, the step where the apiserver may change or reject an object (15.5); a pod's securityContext: runAsUser, runAsNonRoot, fsGroup, capabilities, allowPrivilegeEscalation, seccompProfile (17.37); Pod Security Admission and its restricted level (17.39); capabilities (11.29); ClusterRole / RoleBinding and auth can-i (17.30, 17.35); projects and their UID range annotation (30.4).

The words you need first

An admission plugin with opinions

A SecurityContextConstraints object (SCC, API group security.openshift.io, cluster-scoped) describes what a pod is allowed to ask for and what it gets when it does not ask. Every pod OpenShift admits passes through the SCC admission plugin, which:

  1. collects the SCCs the pod may use (next lesson: who may use what),
  2. tries them in order,
  3. for each one fills in the fields the pod left empty (UID, fsGroup, SELinux level, seccomp, dropped capabilities...) and then validates the result,
  4. admits the pod with the first SCC that validates, and writes its name into the pod's openshift.io/scc annotation - or rejects the pod with every SCC's reason.

Step 3 is the part Kubernetes has no equivalent for. Pod Security Admission only judges a pod; SCC admission also rewrites it. That is why the same YAML behaves differently here.

The SCCs a 4.21 cluster ships

oc get scc lists them. The table is wide; the next paragraph explains the columns you care about, so do not try to read every cell yet:

# cluster-scoped: as kubeadmin
oc get scc
NAME               PRIV    CAPS                   SELINUX     RUNASUSER          FSGROUP     SUPGROUP    PRIORITY     READONLYROOTFS   VOLUMES
anyuid             false   <no value>             MustRunAs   RunAsAny           RunAsAny    RunAsAny    10           false            ["configMap","csi","downwardAPI","emptyDir","ephemeral","persistentVolumeClaim","projected","secret"]
hostaccess         false   <no value>             MustRunAs   MustRunAsRange     MustRunAs   RunAsAny    <no value>   false            ["configMap","csi","downwardAPI","emptyDir","ephemeral","hostPath","persistentVolumeClaim","projected","secret"]
hostmount-anyuid   false   <no value>             MustRunAs   RunAsAny           RunAsAny    RunAsAny    <no value>   false            ["configMap","csi","downwardAPI","emptyDir","ephemeral","hostPath","nfs","persistentVolumeClaim","projected","secret"]
hostnetwork        false   <no value>             MustRunAs   MustRunAsRange     MustRunAs   MustRunAs   <no value>   false            [...]
hostnetwork-v2     false   ["NET_BIND_SERVICE"]   MustRunAs   MustRunAsRange     MustRunAs   MustRunAs   <no value>   false            [...]
nested-container   false   <no value>             MustRunAs   MustRunAsRange     MustRunAs   RunAsAny    <no value>   false            [...]
node-exporter      true    <no value>             RunAsAny    RunAsAny           RunAsAny    RunAsAny    <no value>   false            ["*"]
nonroot            false   <no value>             MustRunAs   MustRunAsNonRoot   RunAsAny    RunAsAny    <no value>   false            [...]
nonroot-v2         false   ["NET_BIND_SERVICE"]   MustRunAs   MustRunAsNonRoot   RunAsAny    RunAsAny    <no value>   false            [...]
privileged         true    ["*"]                  RunAsAny    RunAsAny           RunAsAny    RunAsAny    <no value>   false            ["*"]
restricted         false   <no value>             MustRunAs   MustRunAsRange     MustRunAs   RunAsAny    <no value>   false            [...]
restricted-v2      false   ["NET_BIND_SERVICE"]   MustRunAs   MustRunAsRange     MustRunAs   RunAsAny    <no value>   false            [...]
restricted-v3      false   ["NET_BIND_SERVICE"]   MustRunAs   MustRunAsRange     MustRunAs   RunAsAny    <no value>   false            [...]

(Listing SCCs needs cluster rights; as developer you get Error from server (Forbidden): securitycontextconstraints.security.openshift.io is forbidden.)

Read the columns as "strategies":

What each one is for, in one line:

SCCfor
restricted-v2every authenticated user and ServiceAccount; the default for normal pods
restricted-v34.20+: like restricted-v2 but the pod must run in a user namespace (hostUsers: false: UID 0 inside the container maps to an unprivileged UID on the node)
nonroot-v2a fixed non-root UID of your choosing (runAsUser: 1001) instead of the project's
anyuidany UID, root included. Priority 10. Granted to cluster admins
hostnetwork-v2hostNetwork / hostPort, still non-root
hostmount-anyuid, hostaccesshost paths / host namespaces - node agents
privilegedeverything. The platform's own daemons; nobody else
restricted, nonroot, hostnetworkthe pre-4.11 versions; on fresh 4.11+ installs restricted is not granted to users any more
nested-containerrunning podman/buildah (Docker-like tools that need no daemon) inside a pod

restricted-v3: whether ordinary users are granted it depends on the version and how the cluster was upgraded - check yours with oc describe clusterrolebinding system:openshift:scc:restricted-v3 (or oc adm policy who-can use scc restricted-v3). Either way, a pod that does not set hostUsers: false fails its check and lands on restricted-v2, which is what you will see on almost every pod.

restricted-v2 in detail

oc describe scc NAME prints one SCC as a readable list. Most lines are "no" to something (no privileged, no host network, no host ports); the four Strategy lines at the bottom are the ones that change your pod:

# as kubeadmin
oc describe scc restricted-v2
Name:                                           restricted-v2
Priority:                                       <none>
Access:
  Users:                                        <none>
  Groups:                                       <none>
Settings:
  Allow Privileged:                             false
  Allow Privilege Escalation:                   false
  Default Add Capabilities:                     <none>
  Required Drop Capabilities:                   ALL
  Allowed Capabilities:                         NET_BIND_SERVICE
  Allowed Seccomp Profiles:                     runtime/default
  Allowed Volume Types:                         configMap,csi,downwardAPI,emptyDir,ephemeral,persistentVolumeClaim,projected,secret
  Allowed Flexvolumes:                          <all>
  Allowed Unsafe Sysctls:                       <none>
  Forbidden Sysctls:                            <none>
  Allow Host Network:                           false
  Allow Host Ports:                             false
  Allow Host PID:                               false
  Allow Host IPC:                               false
  Read Only Root Filesystem:                    false
  Run As User Strategy: MustRunAsRange
    UID:                                        <none>
    UID Range Min:                              <none>
    UID Range Max:                              <none>
  SELinux Context Strategy: MustRunAs
    User:                                       <none>
    Role:                                       <none>
    Type:                                       <none>
    Level:                                      <none>
  FSGroup Strategy: MustRunAs
    Ranges:                                     <none>
  Supplemental Groups Strategy: RunAsAny
    Ranges:                                     <none>

"Users: <none>, Groups: <none>" does not mean nobody can use it. oc describe scc only shows the SCC's own users/groups fields. Access is granted through RBAC: the ClusterRole system:openshift:scc:restricted-v2 (verb use on that SCC) is bound to the group system:authenticated. The "UID Range Min/Max: <none>" means "take it from the namespace annotation".

Before and after

You apply a Deployment with no securityContext at all:

      containers:
      - name: web
        image: docker.io/nginxinc/nginx-unprivileged:1.27

Creating it already prints a warning - Pod Security Admission, not SCC:

Warning: would violate PodSecurity "restricted:latest": allowPrivilegeEscalation != false (container "web" must set securityContext.allowPrivilegeEscalation=false), unrestricted capabilities (container "web" must set securityContext.capabilities.drop=["ALL"]), runAsNonRoot != true (pod or container "web" must set securityContext.runAsNonRoot=true), seccompProfile (pod or container "web" must set securityContext.seccompProfile.type to "RuntimeDefault" or "Localhost")
deployment.apps/web created

The project's PSA warn label is restricted (30.4), and PSA looks at the Deployment's template, which really does lack those fields. The warning is harmless noise: the pods themselves will be fixed up by SCC admission before PSA sees them. Teams that want a quiet oc apply put those four fields in their templates - which is also exactly what makes the manifest portable to a strict PSA cluster.

The pod that actually runs (grep -B2 -A12 shows 2 lines before and 12 after each match, 7.1):

# web = the deployment of the next mission
oc get pod -l app=web -o yaml | grep -B2 -A12 'securityContext:'
    securityContext:
      allowPrivilegeEscalation: false
      capabilities:
        drop:
        - ALL
      runAsNonRoot: true
      runAsUser: 1000650000
...
  securityContext:
    fsGroup: 1000650000
    seLinuxOptions:
      level: s0:c27,c14
    seccompProfile:
      type: RuntimeDefault

Everything there was written by admission: the first UID and group of the project range, the MCS level, drop ALL, no escalation, the RuntimeDefault seccomp profile. And admission leaves its signature in an annotation. The jsonpath (15.38) needs \. because the annotation name itself contains dots; oc rsh deploy/web id runs id in one of the Deployment's pods:

oc get pod -l app=web -o jsonpath='{.items[0].metadata.annotations.openshift\.io/scc}{"\n"}'
restricted-v2
oc rsh deploy/web id
uid=1000650000(1000650000) gid=0(root) groups=0(root),1000650000

(the simulator prints uid=1000650000 gid=0(root)...: there is no passwd entry for that UID, so real id also shows just the number unless the image's entrypoint adds one.)

Three facts to take away from that one line:

  1. The UID is not the image's USER. Whatever USER the Dockerfile set - root, 101, 1001 - restricted-v2 replaces it with the project's UID. Different project, different UID; you cannot know it at build time.
  2. The primary group is 0 (root). Not root user - root group. This is the hook that makes arbitrary UIDs workable: an image that makes its writable directories owned by group 0 and group-writable works for any UID.
  3. fsGroup is also added (groups=...,1000650000), so volumes that support ownership management (emptyDir, most block PVCs) are made writable for the pod.

Pod Security Admission is still there

OpenShift runs PSA too, independently. Globally it enforces privileged (nothing blocked) and warns/audits restricted. On your projects a controller syncs the warn and audit labels to match the most privileged SCC your ServiceAccounts can use (restricted-v2 -> restricted, anyuid -> baseline, privileged -> privileged), so that you are not warned about things SCC will let through anyway. Setting a PSA label by hand switches the sync off for that label; security.openshift.io/scc.podSecurityLabelSync=false switches it off for the namespace. Enforcement on OpenShift is SCC's job.

The questions to ask of a pod on OpenShift

  1. Which SCC admitted it? (openshift.io/scc)
  2. Which UID does it actually run as? (oc rsh ... id, or the runAsUser admission wrote)
  3. Does the image write outside a mounted volume, need root, chown files, or bind a port below 1024? If yes, restricted-v2 will break it - lesson 30.9.

What you can now do

Why it helps

This lesson explains the single biggest source of "it works everywhere but OpenShift" tickets. When you know that SCC admission rewrites pods, not only validates them, you stop being surprised that the UID is not the image's USER and differs per project, and you understand why group 0 is the hook that makes arbitrary UIDs workable.

It also clears up the noise: the Pod Security Admission (PSA, upstream Kubernetes' own pod check) warning on oc apply looks alarming but is harmless, because PSA judges the Deployment template while SCC fixes the pods. Being able to read oc get scc columns and explain restricted-v2 versus nonroot-v2 versus anyuid is what platform security reviews and interviews expect, and it is the foundation for the fixes in the next lesson.

FAQ

What is the difference between SCC and Pod Security Admission?

Pod Security Admission, upstream Kubernetes, only validates pods against baseline or restricted levels and warns, audits or rejects. SCC admission, OpenShift's own, selects an SCC for each pod, fills in fields the pod left empty such as UID, fsGroup, SELinux level, seccomp and dropped capabilities, then validates. OpenShift runs both: SCC enforces, while PSA globally only warns and audits, with labels synced from the SCCs your ServiceAccounts can use.

Why does my pod run as UID 1000650000 instead of the image's USER?

Because restricted-v2 uses the MustRunAsRange strategy: pods get the first UID of the project's range from the openshift.io/sa.scc.uid-range annotation, overriding the image's USER. Each project has a different range, so the UID cannot be known when the image is built. Images must therefore work for any UID, typically by making writable paths owned by group 0 and group-writable.

Why is the pod's group 0? Isn't that root?

It is the root group, not the root user. restricted-v2 runs containers with primary group 0 on purpose: an image whose writable directories are owned by group 0 with group write permission works for any UID the project assigns. Group 0 membership grants no special kernel privileges; the user is still non-root, capabilities are dropped and privilege escalation is disabled.

Why does oc apply warn "would violate PodSecurity restricted" when the pod runs fine?

The project's PSA warn label is restricted, and PSA checks the Deployment's pod template, which lacks fields like allowPrivilegeEscalation: false, capabilities.drop: [ALL], runAsNonRoot and seccompProfile. SCC admission adds those to the actual pods afterwards, so they run compliant. To silence the warning and make manifests portable to strict clusters, set those four fields in your templates.

Why does oc describe scc show no users or groups?

oc describe scc shows only the SCC's own users and groups fields. On current OpenShift, access is normally granted through RBAC instead: a ClusterRole such as system:openshift:scc:restricted-v2 with the verb use on that SCC, bound to system:authenticated. Check who can use an SCC with oc adm policy who-can use scc restricted-v2 or oc auth can-i use scc/anyuid --as=....

In an interview Mid

What are SecurityContextConstraints in OpenShift, and what does restricted-v2 do to a pod?

An SCC is a cluster-scoped object listing what a pod may ask for (root, host network, volume types, capabilities) and what it gets when it does not ask. The SCC admission plugin collects the SCCs the pod may use, tries them in order, fills in the empty fields and then validates, admits with the first that fits, and records it in the pod's openshift.io/scc annotation. Pod Security Admission only judges a pod; SCC also rewrites it.

restricted-v2, the default for every authenticated user and ServiceAccount:

So oc rsh deploy/web id prints uid=1000650000 gid=0(root). Group 0 is the hook: images whose writable directories are group-0-writable work for any UID. PSA still runs (warn/audit synced from the SCCs), but enforcement is SCC's.

Also asked: How do you make a container image run correctly under restricted-v2? · How do you find out which SCC admitted a pod and which UID it runs as? · Compare SCCs with Kubernetes Pod Security Admission.

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.