OnCallReady

Lesson 17.39 · Kubernetes: Scheduling, Health & Security · 13 min read

Pod Security Standards and Pod Security Admission

In plain words

Think of a swimming pool with three areas and a sign at each entrance. The toddler pool (restricted) has strict rules: armbands on, no diving, no running. The main pool (baseline) just forbids the really dangerous stuff, like diving into the shallow end. The staff-only area (privileged) has no rules at all. A lifeguard at each entrance checks swimmers against that area's sign. The lifeguard can turn people away (enforce), just warn them (warn), or write their name in a logbook (audit).

That's Pod Security Admission: each namespace gets labels like pod-security.kubernetes.io/enforce: restricted, and the API server checks every new pod against that level of the Pod Security Standards. One catch: enforce checks pods, not Deployments, so a bad Deployment is accepted and its pods quietly fail.

Three levels, three modes, one label each

The problem. You hardened one Deployment by hand (17.37). The next team ships a root, privileged pod anyway. The cluster should refuse unsafe pods by itself, per namespace - and you need to switch that on without breaking what already runs.

What you need to know already: securityContext (17.37), admission (17.9), namespace labels (15.26), ReplicaSet FailedCreate events (17.9).

PodSecurityPolicy (an older mechanism) was removed in 1.25. Its replacement is built into the apiserver: Pod Security Admission (PSA) checks pods against the Pod Security Standards, configured per namespace with labels:

apiVersion: v1
kind: Namespace
metadata:
  name: payments
  labels:
    pod-security.kubernetes.io/enforce: restricted      # reject violating pods
    pod-security.kubernetes.io/enforce-version: v1.34   # pin the rules (default: latest)
    pod-security.kubernetes.io/warn: restricted         # warn the client (kubectl prints it)
    pod-security.kubernetes.io/audit: restricted        # annotate the audit log

Levels:

Modes: enforce rejects, warn returns a warning to the client, audit records it in the apiserver audit log (a file where the apiserver can record every request: who did what, when). They can be set to different levels - the standard migration is enforce: baseline + warn: restricted while teams fix their manifests.

What a rejection looks like

$ kubectl run web --image=nginx:1.27 -n payments
Error from server (Forbidden): pods "web" is forbidden: violates PodSecurity "restricted:latest": allowPrivilegeEscalation != false (container "web" must set securityContext.allowPrivilegeEscalation=false), unrestricted capabilities (container "web" must set securityContext.capabilities.drop=["ALL"]), runAsNonRoot != true (pod or container "web" must set securityContext.runAsNonRoot=true), seccompProfile (pod or container "web" must set securityContext.seccompProfile.type to "RuntimeDefault" or "Localhost")

The message is a checklist: each reason (detail) names the field to set. The four above are the ones every unmodified image fails under restricted.

Baseline failures read the same way:

pods "debug" is forbidden: violates PodSecurity "baseline:latest": host namespaces (hostNetwork=true), hostPath volumes (volume "root"), privileged (container "debug" must not set securityContext.privileged=true)

The workload trap: enforce only rejects PODS

PSA's enforce mode is checked on pods. A Deployment is not a pod, so it is accepted - and then its ReplicaSet cannot create a single pod:

# an illustration: the PSS mission's namespaces and web.yaml
kubectl apply -f web.yaml -n payments
deployment.apps/web created                       <- no error!
kubectl get deploy web -n payments
NAME   READY   UP-TO-DATE   AVAILABLE   AGE
web    0/3     0            0           2m
kubectl describe rs -n payments -l app=web | tail -2
  Warning  FailedCreate  5s (x14 over 2m)  replicaset-controller  Error creating: pods "web-7c9f4d8b6c-" is forbidden: violates PodSecurity "restricted:latest": allowPrivilegeEscalation != false (...), ...

That is why you set warn as well: warn and audit are evaluated on workload objects (Deployments, StatefulSets, Jobs, CronJobs... - anything with a pod template), so the person applying the Deployment sees it immediately:

# an illustration: the PSS mission's namespaces and web.yaml
kubectl apply -f web.yaml -n payments
Warning: would violate PodSecurity "restricted:latest": allowPrivilegeEscalation != false (container "web" must set securityContext.allowPrivilegeEscalation=false), unrestricted capabilities (container "web" must set securityContext.capabilities.drop=["ALL"]), runAsNonRoot != true (pod or container "web" must set securityContext.runAsNonRoot=true), seccompProfile (pod or container "web" must set securityContext.seccompProfile.type to "RuntimeDefault" or "Localhost")
deployment.apps/web created

Labelling a namespace that already has pods

PSA never evicts running pods. When you add or tighten enforce, it checks the existing pods and warns:

# an illustration: the PSS mission's namespaces and web.yaml
kubectl label ns shop pod-security.kubernetes.io/enforce=restricted
Warning: existing pods in namespace "shop" violate the new PodSecurity enforce level "restricted:latest"
Warning: web-5d8f7c9b4d-2kq9x (and 2 other pods): allowPrivilegeEscalation != false, unrestricted capabilities, runAsNonRoot != true, seccompProfile
namespace/shop labeled

Those pods keep running - until they are replaced (a rollout, a node drain, a crash), and then the replacement is rejected. To preview without changing anything:

# an illustration: the PSS mission's namespaces and web.yaml
kubectl label ns shop pod-security.kubernetes.io/enforce=restricted --dry-run=server
Warning: existing pods in namespace "shop" violate the new PodSecurity enforce level "restricted:latest"
...
namespace/shop labeled (server dry run)

A misspelled level is refused by the namespace validation (must be one of privileged, baseline, restricted) - but a misspelled label key is just an unknown label and silently does nothing. Check with kubectl get ns shop --show-labels.

Versions and exemptions

enforce-version: v1.34 pins which version of the standard is checked, so a cluster upgrade cannot suddenly start rejecting pods because the standard got stricter; latest follows the cluster. Cluster-wide defaults and exemptions (namespaces, usernames, runtime classes) are set in the apiserver's AdmissionConfiguration for the PodSecurity plugin - that is how kube-system (calico needs hostNetwork and privileged) stays privileged while everything else defaults to baseline.

PSA is deliberately simple: three fixed levels, per namespace. For anything more specific ("images only from our registry", "every pod must have resource limits") teams add a policy engine - an admission add-on that checks your own rules - such as Kyverno, OPA Gatekeeper, or the built-in ValidatingAdmissionPolicy (rules written in CEL, a small expression language) - on top.

What you can now do

Why it helps

This is how a platform enforces the securityContext rules from the last lesson across hundreds of namespaces, and you'll own the labels. The typical rollout, enforce: baseline with warn: restricted, lets teams see what they need to fix before it's enforced.

The support tickets are predictable. "Our Deployment applied fine but has 0 pods": enforce rejected the pods, and the reason is on the ReplicaSet. "Everything broke after the node drain": the namespace was labelled restricted weeks ago, existing pods kept running, and the replacements are rejected. You'll know to use --dry-run=server before labelling, to set warn alongside enforce, and when to reach for Kyverno or ValidatingAdmissionPolicy for rules PSA can't express.

Commands in this lesson

kubectl

FAQ

What replaced PodSecurityPolicy?

Pod Security Admission, built into the API server. PodSecurityPolicy was removed in Kubernetes 1.25. PSA is simpler: three fixed levels (privileged, baseline, restricted) from the Pod Security Standards, applied per namespace with labels, in three modes (enforce, warn, audit). For custom rules you add a policy engine like Kyverno, OPA Gatekeeper or ValidatingAdmissionPolicy.

My Deployment was created but has 0 pods. Is PSA involved?

Likely. enforce is evaluated on pods, not on Deployments, so the Deployment is accepted and then the ReplicaSet controller can't create a single pod. The rejection is in k describe rs as FailedCreate ... violates PodSecurity "restricted:latest". Setting warn on the namespace shows the problem at apply time, because warn and audit also check workload templates.

Does labelling a namespace restricted kill running pods?

No, PSA never evicts. When you add or tighten enforce, it warns about existing violating pods and they keep running. The problem appears when they're replaced by a rollout, a drain or a crash: the replacements are rejected. Preview first with kubectl label ns <ns> pod-security.kubernetes.io/enforce=restricted --dry-run=server.

What does enforce-version do?

It pins which version of the standard is checked, for example v1.34. Without it, latest follows the cluster version, so a cluster upgrade that makes a level stricter could suddenly reject pods that were fine before. Pinning lets you upgrade the cluster and the policy separately.

What happens if I misspell a label?

A misspelled level value (restricted spelled wrong) is rejected by namespace validation, which lists the valid values. A misspelled label key is just an unknown label and silently does nothing: no enforcement, no warning. Check with kubectl get ns <ns> --show-labels and test with a pod that should be rejected.

In an interview Mid

What are the Pod Security Standards, and how would you roll out restricted without breaking teams?

Pod Security Admission is built into the API server and checks pods against three levels, set per namespace with labels:

Three modes: enforce rejects, warn tells the client, audit logs. Roll out in steps:

  1. enforce: baseline + warn: restricted (and audit: restricted) - nothing breaks, everyone sees what would.
  2. Preview: kubectl label ns shop pod-security.kubernetes.io/enforce=restricted --dry-run=server lists the violating pods.
  3. Fix the manifests, then enforce restricted; pin enforce-version so upgrades do not change the rules.

The trap: enforce checks pods, so a Deployment is accepted and its ReplicaSet then cannot create pods - which is why warn (checked on workloads) matters. Running pods are never evicted, only their replacements rejected.

Also asked: What are the limits of Pod Security Admission, and what would you add on top? · Why does a Deployment get accepted in a restricted namespace and still create no pods? · What is the difference between the enforce, warn and audit modes?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.