OnCallReady

Lesson 30.9 · OpenShift · 26 min read

SecurityContextConstraints II: selection, rejections, and fixing images

In plain words

Imagine a theme park where each ride has a height rule, and a kid's wristband lists which rides they may try. At the gate, staff try the rides on the wristband in a fixed order, strictest first, and send the kid to the first ride they fit. If none fit, they write down, ride by ride, why not: "too short for ride A, ride B not on your wristband". And a parent buying tickets in person gets a different wristband than a kid sent in by the school bus.

SCC selection works like that. Admission collects the SCCs the pod's ServiceAccount may use (plus the creating user's, only for pods created directly), sorts them by priority then restrictiveness, and admits with the first that validates. A rejection lists every SCC with its reason. oc adm policy scc-subject-review asks the question without creating anything.

Why "just give it anyuid" is the wrong fix

Your Helm chart (25.19) from AKS creates a Deployment, and no pods appear - or pods appear and crash with Permission denied. The CLI itself suggests one command that makes it go away. This lesson is how to read what went wrong, fix the image or the manifest, and - when a special grant really is needed - make it as small as possible.

What you need to know already: SCCs, strategies and restricted-v2 (30.7); how a Deployment's pods are created by a ReplicaSet (15.16); ServiceAccounts and serviceAccountName (17.33); Role, RoleBinding, ClusterRoleBinding (17.30); a Dockerfile's USER, RUN and file ownership (10.8, 10.40); chgrp / chmod and group permissions (4.5); emptyDir volumes (16.37). New here: ports below 1024 need root or the NET_BIND_SERVICE capability.

The words you need first

Who may use which SCC

For every pod, admission builds the set of SCCs usable by:

The second point bites. A Deployment's pods are created by the ReplicaSet controller, not by you, so for them only the ServiceAccount counts. A cluster admin who oc runs a test pod gets anyuid (admins can use it and its priority is 10)

is a classic.

"Usable" means either the SCC's own users/groups fields name the subject, or RBAC grants the verb use on it. --as= asks can-i on behalf of someone else (17.35); oc adm policy who-can VERB RESOURCE lists everyone who may:

$ oc auth can-i use scc/anyuid --as=system:serviceaccount:shop:default -n shop
no
$ oc adm policy who-can use scc anyuid -n shop
resourceaccessreviewresponse.authorization.openshift.io/<unknown>

Namespace: shop
Verb:      use
Resource:  securitycontextconstraints.security.openshift.io

Users:  kube:admin
        system:admin
Groups: system:cluster-admins
        system:masters

The order

The usable SCCs are sorted:

  1. by priority, highest first (unset = 0),
  2. then from most restrictive to least restrictive (a points system: privileged, host network/ports, host volumes, RunAsAny... add points),
  3. then by name.

and admission takes the first one the pod validates against. That is why a normal pod lands on restricted-v2 even when it could also use nonroot-v2 or anyuid: it is tried first and it fits. Priority is the dangerous knob - an SCC with a high priority and a broad grant will "steal" pods that restricted-v2 would have admitted, and pods quietly start running as root. Pin the one you mean with an annotation on the pod template (the annotations under spec.template.metadata end up on every pod, 15.26):

spec:
  template:
    metadata:
      annotations:
        openshift.io/required-scc: nonroot-v2

With openshift.io/required-scc, only that SCC is tried; if it does not fit (or the ServiceAccount may not use it) the pod is rejected instead of silently falling to another one.

Reading the rejection

A chart that worked on AKS, with the pod pinned to root:

      securityContext:
        runAsUser: 0
        fsGroup: 0

oc get pods shows nothing - there is no pod to show. The Deployment is created, the ReplicaSet exists, and its events carry the reason. oc describe rs -l app=legacy describes the ReplicaSets with that label; tail -3 keeps the last three lines, where the events are. Do not try to read the long line yet - the paragraph after it takes it apart:

# the legacy/reporting scenario of the next missions (the adm commands as kubeadmin)
oc get deploy legacy
NAME     READY   UP-TO-DATE   AVAILABLE   AGE
legacy   0/1     0            0           40s
oc describe rs -l app=legacy | tail -3
  Type     Reason        Age               From                   Message
  ----     ------        ----              ----                   -------
  Warning  FailedCreate  4s (x5 over 40s)  replicaset-controller  Error creating: pods "legacy-v9g8hnvcvt-" is forbidden: unable to validate against any security context constraint: [provider "anyuid": Forbidden: not usable by user or serviceaccount, provider restricted-v2: .spec.securityContext.fsGroup: Invalid value: []int64{0}: 0 is not an allowed group, spec.containers[0].securityContext.runAsUser: Invalid value: 0: must be in the ranges: [1000650000, 1000659999], provider "restricted-v3": Forbidden: not usable by user or serviceaccount, provider "restricted": Forbidden: not usable by user or serviceaccount, provider "nested-container": Forbidden: not usable by user or serviceaccount, provider "nonroot-v2": Forbidden: not usable by user or serviceaccount, provider "nonroot": Forbidden: not usable by user or serviceaccount, provider "hostmount-anyuid": Forbidden: not usable by user or serviceaccount, provider "hostnetwork-v2": Forbidden: not usable by user or serviceaccount, provider "hostnetwork": Forbidden: not usable by user or serviceaccount, provider "hostaccess": Forbidden: not usable by user or serviceaccount, provider "node-exporter": Forbidden: not usable by user or serviceaccount, provider "privileged": Forbidden: not usable by user or serviceaccount]

How to read that wall: every SCC in the cluster appears once, in the order admission tried them. Skip every Forbidden: not usable by user or serviceaccount - those were never options. The providers without quotes are the ones the pod could use, and the text after the colon is exactly why each failed. Here only restricted-v2 was usable, and it said: fsGroup 0 is not in the range, runAsUser 0 is not in the range. The fix is in the manifest: delete both lines (let admission pick), or put values inside the range - which you cannot know in advance, so: delete them.

Other rejections you will meet (hostNetwork: the pod uses the node's own network interfaces; hostPort: a port opened on the node itself; NET_ADMIN: the capability to change network settings):

provider textthe pod asked for
spec.containers[0].securityContext.privileged: Invalid value: true: Privileged containers are not allowedprivileged: true
spec.securityContext.hostNetwork: Invalid value: true: Host network is not allowed to be usedhostNetwork: true
spec.volumes[0]: Invalid value: "hostPath": hostPath volumes are not allowed to be useda hostPath volume
spec.containers[0].securityContext.capabilities.add: Invalid value: "NET_ADMIN": capability may not be addedcapabilities.add: [NET_ADMIN]
spec.containers[0].ports[0].hostPort: Invalid value: 8080: Host ports are not allowed to be useda hostPort
.spec.hostUsers: Invalid value: true: must be false (restricted-v3)nothing - it just is not a user-namespace pod

Why the Docker Hub image fails

Most crash loops on OpenShift are not rejections - the pod is admitted by restricted-v2 and then the process fails as an unknown non-root UID. The official nginx image is the textbook case. Its log, line by line: the startup script tries to edit a config file and cannot (a warning), nginx notices it is not root (a warning), then it cannot create its cache directory ([emerg] = fatal, 13 = the EACCES error number):

oc logs deploy/legacy
/docker-entrypoint.sh: /docker-entrypoint.d/ is not empty, will attempt to perform configuration
...
10-listen-on-ipv6-by-default.sh: info: can not modify /etc/nginx/conf.d/default.conf (read-only file system?)
...
/docker-entrypoint.sh: Configuration complete; ready for start up
nginx: [warn] the "user" directive makes sense only if the master process runs with super-user privileges, ignored in /etc/nginx/nginx.conf:2
nginx: [emerg] mkdir() "/var/cache/nginx/client_temp" failed (13: Permission denied)

/var/cache/nginx is owned by root in the image; UID 1000650000 cannot create directories there. Give it an emptyDir at /var/cache/nginx and /var/run and it gets one step further - and fails on the next root assumption:

nginx: [emerg] bind() to 0.0.0.0:80 failed (13: Permission denied)

Ports below 1024 need root or NET_BIND_SERVICE, which restricted-v2 drops. The list of root assumptions to look for in any image:

  1. USER missing or USER root - fine on its own under restricted-v2 (you get the project UID anyway), but everything below usually follows from it;
  2. writes outside a volume: /var/cache, /var/run, /var/log/<app>, the app directory itself;
  3. chown/chmod at startup (Operation not permitted);
  4. binding port 80/443;
  5. fsGroup: 0 / runAsUser: 0 hard-coded in the chart.

Fix the image, not the SCC

The portable image - it runs on OpenShift, on AKS, on a strict PSA cluster. UBI (Universal Base Image) is Red Hat's freely redistributable base image family, the usual FROM on OpenShift; its images already follow the rules below:

FROM registry.access.redhat.com/ubi9/nginx-126
# or FROM docker.io/nginxinc/nginx-unprivileged:1.27 - listens on 8080, writes to /tmp
COPY html/ /opt/app-root/src/
EXPOSE 8080
USER 1001

and for your own applications, the Red Hat guideline for arbitrary UIDs:

RUN mkdir -p /app/data /app/cache && \
    chgrp -R 0 /app && \
    chmod -R g=u /app
USER 1001

chgrp -R 0 makes the directories owned by the root group; chmod g=u gives the group the same rights as the owner. Every pod on OpenShift runs with group 0, so whatever the UID turns out to be, it can write there. USER 1001 keeps the image non-root everywhere else.

In the Kubernetes objects: listen on 8080 (Service port: 80, targetPort: 8080 if you want to keep the Service port), mount emptyDir volumes for scratch directories, and do not hard-code runAsUser/fsGroup.

When the need is real

Some images genuinely need a fixed UID - a vendor product whose data files are owned by 1001, a database that insists on its own user. The narrowest grant is nonroot-v2 (any fixed non-root UID, everything else still restricted), to a dedicated ServiceAccount, in one namespace:

oc create sa makes the ServiceAccount; oc adm policy add-scc-to-user SCC -z SA allows that ServiceAccount to use the SCC (it creates a RoleBinding to the SCC's use ClusterRole):

oc create sa reporting -n shop
oc adm policy add-scc-to-user nonroot-v2 -z reporting -n shop
clusterrole.rbac.authorization.k8s.io/system:openshift:scc:nonroot-v2 added: "reporting"
oc get rolebinding system:openshift:scc:nonroot-v2 -n shop -o wide
NAME                              ROLE                                          AGE   USERS   GROUPS   SERVICEACCOUNTS
system:openshift:scc:nonroot-v2   ClusterRole/system:openshift:scc:nonroot-v2   5s                     shop/reporting

then in the Deployment: serviceAccountName: reporting, runAsUser: 1001, and openshift.io/required-scc: nonroot-v2 on the template.

Watch the -n: oc adm policy add-scc-to-user with an explicit -n creates a RoleBinding in that namespace. Without -n it uses your current project for the ServiceAccount but binds with a ClusterRoleBinding - valid everywhere. (Same command, very different blast radius. Check with oc get clusterrolebinding system:openshift:scc:nonroot-v2.) The GitOps-friendly equivalent is a Role:

$ oc create role use-nonroot-v2 --verb=use --resource=scc --resource-name=nonroot-v2 -n shop
$ oc create rolebinding reporting-nonroot-v2 --role=use-nonroot-v2 --serviceaccount=shop:reporting -n shop

Why not anyuid

oc status summarises a project; --suggest adds advice for each problem it finds:

oc status --suggest
...
Errors:
  * pod/legacy-5d6b9c6f8-abcde is crash-looping

    The container is starting and exiting repeatedly. This usually means the container is unable
    to start, misconfigured, or limited by security restrictions. Check the container logs with

      oc logs legacy-5d6b9c6f8-abcde -c legacy

    Current security policy prevents your containers from being run as the root user. Some images
    may fail expecting to be able to change ownership or permissions on directories. Your admin
    can grant you access to run containers that need to run as the root user with this command:

      oc adm policy add-scc-to-user anyuid -n shop -z default

The CLI itself suggests it, and it "works". What it does: every pod using the default ServiceAccount in that project may now run as UID 0, with the capabilities restricted-v2 dropped (anyuid only drops MKNOD) and privilege escalation allowed - and because anyuid has priority 10, pods that would have been fine on restricted-v2 now land on anyuid too. A container escape from a root container is a root process on the node. In a bank, anyuid on default fails every security review, and it hides the real bug: the image is not portable and will fail the next strict cluster too.

The review commands

oc adm policy scc-subject-review -f legacy.yaml
RESOURCE            ALLOWED BY
Deployment/legacy   <none>
oc adm policy scc-subject-review -z reporting -f reporting.yaml
RESOURCE               ALLOWED BY
Deployment/reporting   nonroot-v2
oc adm policy scc-review -f legacy.yaml
RESOURCE            SERVICE ACCOUNT   ALLOWED BY
Deployment/legacy   default           anyuid

Both run admission as a dry run: nothing is created. They are the fastest way to answer "why does this manifest not get pods" without reading the 2 KB event. (-f FILE = the manifest to check.)

What you can now do

Why it helps

This is the lesson for the first incident: a chart that worked on AKS gets zero pods, and the ReplicaSet event is a 2 KB wall of text. Knowing to skip every not usable by user or serviceaccount and read the unquoted providers turns that wall into "runAsUser 0 is not in the range, delete that line" in a minute.

It also stops you from making the classic mistake. oc status --suggest literally recommends oc adm policy add-scc-to-user anyuid -z default; doing it lets every pod in the project run as root, and anyuid's priority 10 pulls in pods that were fine before. In a bank that fails security review. The right fix is the image, or at most nonroot-v2 on a dedicated ServiceAccount in one namespace, with -n so it is a RoleBinding and not a cluster-wide grant.

Commands in this lesson

oc

FAQ

Why does my pod work with oc run but not in a Deployment?

Because for directly created pods, SCC admission considers both the creating user's and the ServiceAccount's SCCs, while Deployment pods are created by the ReplicaSet controller, so only the ServiceAccount counts. A cluster admin running oc run can use anyuid, which has priority 10, while the same spec in a Deployment gets restricted-v2 and may be rejected or crash as a non-root UID.

How do I read a "unable to validate against any security context constraint" error?

It lists every SCC in the order admission tried them. Ignore the ones saying Forbidden: not usable by user or serviceaccount; they were never options. The providers written without quotes were usable, and the text after each colon is the precise reason that SCC rejected the pod, for example runAsUser: Invalid value: 0: must be in the ranges. Fix those fields in the manifest.

Why not just grant anyuid?

anyuid lets pods run as any UID including root, keeps most capabilities (it only drops MKNOD) and allows privilege escalation. Granted to the default ServiceAccount, it applies to every pod in the project, and its priority of 10 means pods that were fine on restricted-v2 now land on anyuid too. A container escape then yields root on the node. It also hides the real problem: the image is not portable.

What does openshift.io/required-scc do?

It is a pod or pod-template annotation that pins admission to exactly one SCC. Only that SCC is tried; if the ServiceAccount cannot use it or the pod does not validate against it, the pod is rejected instead of silently falling to another SCC. It protects against a high-priority SCC granted later "stealing" pods, and makes the intended security posture explicit in the manifest.

What is the difference between add-scc-to-user with and without -n?

With an explicit -n shop, oc adm policy add-scc-to-user nonroot-v2 -z reporting -n shop creates a RoleBinding in that namespace, so the grant applies only there. Without -n, it uses your current project for the ServiceAccount name but creates a ClusterRoleBinding, which grants the SCC across all namespaces. Always pass -n, or better, manage an explicit Role and RoleBinding for use on the SCC in git.

In an interview Mid

A vendor application must run as UID 1001. How do you allow it securely on OpenShift?

First check it really needs that: most images only need arbitrary-UID fixes (group-0-owned writable directories with chmod g=u, a port above 1024, emptyDirs) and then run fine under restricted-v2.

If the fixed UID is real, the narrowest grant:

  1. A dedicated ServiceAccount: oc create sa reporting -n shop.
  2. nonroot-v2 (any fixed non-root UID, everything else still restricted), granted to that SA in one namespace: oc adm policy add-scc-to-user nonroot-v2 -z reporting -n shop - with -n it is a RoleBinding; without it, a ClusterRoleBinding. (Or a Role with verb use on the SCC, for GitOps.)
  3. In the Deployment: serviceAccountName: reporting, runAsUser: 1001, and openshift.io/required-scc: nonroot-v2 on the pod template, so it never silently lands on another SCC.
  4. Verify: oc adm policy scc-subject-review -z reporting -f ..., then the pod's openshift.io/scc annotation.

Remember only the ServiceAccount counts for Deployment pods (the ReplicaSet controller creates them). And not anyuid on default: every pod there could run as UID 0, and with priority 10 it steals pods restricted-v2 would have admitted.

Also asked: How does OpenShift decide which SCC applies to a pod? · Why is granting anyuid to the default ServiceAccount a bad fix? · How do you read an "unable to validate against any security context constraint" event?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.