Three levels, three modes, one label each
The problem. You hardened one Deployment by hand (17.37). The next team ships a root, privileged pod anyway. The cluster should refuse unsafe pods by itself, per namespace - and you need to switch that on without breaking what already runs.
What you need to know already: securityContext (17.37), admission (17.9), namespace labels (15.26), ReplicaSet FailedCreate events (17.9).
PodSecurityPolicy (an older mechanism) was removed in 1.25. Its replacement is built into the apiserver: Pod Security Admission (PSA) checks pods against the Pod Security Standards, configured per namespace with labels:
apiVersion: v1
kind: Namespace
metadata:
name: payments
labels:
pod-security.kubernetes.io/enforce: restricted # reject violating pods
pod-security.kubernetes.io/enforce-version: v1.34 # pin the rules (default: latest)
pod-security.kubernetes.io/warn: restricted # warn the client (kubectl prints it)
pod-security.kubernetes.io/audit: restricted # annotate the audit log
Levels:
- privileged - no restrictions (the default when no label is set).
- baseline - blocks the known privilege escalations:
privileged, host namespaces (hostNetwork/hostPID/hostIPC: sharing the node's own network, process list or IPC instead of the container's private ones), hostPath volumes (a node directory mounted into the pod), hostPorts, capabilities beyond the default set,seccompProfile: Unconfined, unsafe sysctls (kernel settings), AppArmor/SELinux overrides (the two Linux security modules that confine programs), /proc mount changes. - restricted - baseline plus hardening:
runAsNonRoot: true, norunAsUser: 0,allowPrivilegeEscalation: false,capabilities.drop: ["ALL"](only NET_BIND_SERVICE may be added),seccompProfileRuntimeDefault or Localhost, and only these volume types: configMap, csi, downwardAPI, emptyDir, ephemeral, persistentVolumeClaim, projected, secret.
Modes: enforce rejects, warn returns a warning to the client, audit records it in the apiserver audit log (a file where the apiserver can record every request: who did what, when). They can be set to different levels - the standard migration is enforce: baseline + warn: restricted while teams fix their manifests.
What a rejection looks like
$ kubectl run web --image=nginx:1.27 -n payments
Error from server (Forbidden): pods "web" is forbidden: violates PodSecurity "restricted:latest": allowPrivilegeEscalation != false (container "web" must set securityContext.allowPrivilegeEscalation=false), unrestricted capabilities (container "web" must set securityContext.capabilities.drop=["ALL"]), runAsNonRoot != true (pod or container "web" must set securityContext.runAsNonRoot=true), seccompProfile (pod or container "web" must set securityContext.seccompProfile.type to "RuntimeDefault" or "Localhost")
The message is a checklist: each reason (detail) names the field to set. The four above are the ones every unmodified image fails under restricted.
Baseline failures read the same way:
pods "debug" is forbidden: violates PodSecurity "baseline:latest": host namespaces (hostNetwork=true), hostPath volumes (volume "root"), privileged (container "debug" must not set securityContext.privileged=true)
The workload trap: enforce only rejects PODS
PSA's enforce mode is checked on pods. A Deployment is not a pod, so it is accepted - and then its ReplicaSet cannot create a single pod:
# an illustration: the PSS mission's namespaces and web.yaml
kubectl apply -f web.yaml -n payments
deployment.apps/web created <- no error!
kubectl get deploy web -n payments
NAME READY UP-TO-DATE AVAILABLE AGE
web 0/3 0 0 2m
kubectl describe rs -n payments -l app=web | tail -2
Warning FailedCreate 5s (x14 over 2m) replicaset-controller Error creating: pods "web-7c9f4d8b6c-" is forbidden: violates PodSecurity "restricted:latest": allowPrivilegeEscalation != false (...), ...
That is why you set warn as well: warn and audit are evaluated on workload objects (Deployments, StatefulSets, Jobs, CronJobs... - anything with a pod template), so the person applying the Deployment sees it immediately:
# an illustration: the PSS mission's namespaces and web.yaml
kubectl apply -f web.yaml -n payments
Warning: would violate PodSecurity "restricted:latest": allowPrivilegeEscalation != false (container "web" must set securityContext.allowPrivilegeEscalation=false), unrestricted capabilities (container "web" must set securityContext.capabilities.drop=["ALL"]), runAsNonRoot != true (pod or container "web" must set securityContext.runAsNonRoot=true), seccompProfile (pod or container "web" must set securityContext.seccompProfile.type to "RuntimeDefault" or "Localhost")
deployment.apps/web created
Labelling a namespace that already has pods
PSA never evicts running pods. When you add or tighten enforce, it checks the existing pods and warns:
# an illustration: the PSS mission's namespaces and web.yaml
kubectl label ns shop pod-security.kubernetes.io/enforce=restricted
Warning: existing pods in namespace "shop" violate the new PodSecurity enforce level "restricted:latest"
Warning: web-5d8f7c9b4d-2kq9x (and 2 other pods): allowPrivilegeEscalation != false, unrestricted capabilities, runAsNonRoot != true, seccompProfile
namespace/shop labeled
Those pods keep running - until they are replaced (a rollout, a node drain, a crash), and then the replacement is rejected. To preview without changing anything:
# an illustration: the PSS mission's namespaces and web.yaml
kubectl label ns shop pod-security.kubernetes.io/enforce=restricted --dry-run=server
Warning: existing pods in namespace "shop" violate the new PodSecurity enforce level "restricted:latest"
...
namespace/shop labeled (server dry run)
A misspelled level is refused by the namespace validation (must be one of privileged, baseline, restricted) - but a misspelled label key is just an unknown label and silently does nothing. Check with kubectl get ns shop --show-labels.
Versions and exemptions
enforce-version: v1.34 pins which version of the standard is checked, so a cluster upgrade cannot suddenly start rejecting pods because the standard got stricter; latest follows the cluster. Cluster-wide defaults and exemptions (namespaces, usernames, runtime classes) are set in the apiserver's AdmissionConfiguration for the PodSecurity plugin - that is how kube-system (calico needs hostNetwork and privileged) stays privileged while everything else defaults to baseline.
PSA is deliberately simple: three fixed levels, per namespace. For anything more specific ("images only from our registry", "every pod must have resource limits") teams add a policy engine - an admission add-on that checks your own rules - such as Kyverno, OPA Gatekeeper, or the built-in ValidatingAdmissionPolicy (rules written in CEL, a small expression language) - on top.
What you can now do
- Label a namespace with enforce/warn/audit levels and pinned versions.
- Preview a change with
--dry-run=serverand read the rejection checklist. - Migrate safely: enforce baseline + warn restricted, fix, then enforce restricted.