OpenShift: interview questions
The question you are most likely to get for each topic, a model answer, and what else comes up. From chapter 30 of the course.
A Helm chart deploys fine on AKS, but on OpenShift the pods crash-loop or never appear. How do you troubleshoot it? Mid
Almost always SecurityContextConstraints. On OpenShift the SCC admission plugin rewrites every pod: under restricted-v2 it runs as a UID from the project's range (like 1000650000), group 0, all capabilities dropped, no privilege escalation.
No pods at all - the pod was rejected. oc get pods is empty; oc describe rs -l app=... shows unable to validate against any security context constraint. Skip the "not usable" lines; the usable SCC's text names the field - often a hard-coded runAsUser: 0 or fsGroup: 0. Remove them. oc adm policy scc-subject-review -f checks it dry.
Pods crash-loop - admitted, but the process assumed root: Permission denied writing /var/cache/nginx, bind() to 0.0.0.0:80 failed, chown at startup. oc logs, and openshift.io/scc plus oc rsh ... id show which SCC and UID.
Fix the image, not the SCC: a non-root port (8080), writable directories owned by group 0 (chgrp -R 0, chmod -R g=u), emptyDirs for scratch paths, no hard-coded UIDs - or a UBI / nginx-unprivileged base. If a fixed UID is truly needed: nonroot-v2 for a dedicated ServiceAccount in one namespace, never anyuid on default.
Also asked: What is OpenShift and how does it differ from vanilla Kubernetes? · Users get the "Application is not available" page. Walk through your troubleshooting. · How are OpenShift clusters upgraded?
What is OpenShift and how does it differ from vanilla Kubernetes? Mid
OpenShift is Red Hat's Kubernetes distribution: ordinary Kubernetes underneath - same apiserver, Pods, Deployments, Services, RBAC, YAML - plus opinions (defaults you cannot easily turn off) and extra APIs:
- Projects - namespaces requested self-service through a template, each with its own UID range.
- SecurityContextConstraints - every pod runs as a random non-root UID (restricted-v2) unless someone grants more. The reason root images that work on AKS fail here.
- Routes and a built-in HAProxy router (Ingress still works).
- An internal registry, ImageStreams and Builds (S2I) on the cluster.
- Operators everywhere: the platform itself (each reports in
oc get co), and add-ons through OLM / OperatorHub. - Upgrades by the Cluster Version Operator, OS included: nodes run RHCOS, an immutable OS managed by the cluster.
- A built-in OAuth server with
oc login, andoc= kubectl plus OpenShift commands.
The family: OCP (self-installed), OKD (community), ARO / ROSA (managed on Azure / AWS). First commands on an unknown cluster: oc get clusterversion, oc get co, oc get nodes.
Also asked: How do you check the health of an OpenShift cluster? · Explain the operator-based architecture of OpenShift 4. · Why should you never deploy into openshift-* namespaces?
Learn it: 30.1 What OpenShift adds to Kubernetes
How do you log in to an OpenShift cluster from the command line, and what do you check when it fails? Mid
oc login https://api.<cluster>.<domain>:6443 -u <user> - oc finds the cluster's OAuth server, sends your credentials, and gets an access token (sha256~..., 24 h by default). It never stores the password; it writes a cluster, user and context into ~/.kube/config, which kubectl reads too. Or oc login --token=... --server=... with the console's Copy login command.
Where am I: oc whoami, -t (token), --show-server, --show-context.
Failures, read in order:
connection refused/no such host- wrong host or port (the API isapi.<cluster>, port 6443; the console is under*.apps).- an unknown-CA prompt - answer
nand pass--certificate-authority. 401 Unauthorized/token provided is invalid or expired/You must be logged in- who are you? Log in again.403 Forbidden- known user, not allowed: a RoleBinding is missing.
Plus the special identities: kubeadmin (temporary, delete it once a real identity provider exists) and system:admin (the installer's certificate kubeconfig, the break-glass when OAuth is down). oc logout revokes the token on the server.
Also asked: How should a pipeline authenticate to OpenShift? · The cluster's identity provider is down and nobody can log in. What do you do? · What is the difference between a 401 and a 403 from the OpenShift API?
Learn it: 30.2 The oc CLI and logging in
What is an OpenShift project and how does it relate to a Kubernetes namespace? Mid
A project is a namespace seen through a second API (project.openshift.io), with a front door:
- Self-service: any user with the
self-provisionerrole runsoc new-project shop(a ProjectRequest); the API server creates the namespace on their behalf and makes them its admin. Creating a plain namespace needs cluster rights. Many enterprises switch self-provisioning off and hand out projects via tickets or GitOps. - Visibility:
oc get projectslists only the projects you have a role in; others are invisible. - Annotations that matter:
openshift.io/sa.scc.uid-range(the project's block of 10,000 UIDs - restricted-v2 pods run as its first one), supplemental groups (fsGroup), andsa.scc.mcs(the SELinux category isolating it from other projects). - Contents:
builder,deployeranddefaultServiceAccounts, and RoleBindings (adminfor you,system:image-pullers).
Admins shape every new project with a project request template: quotas, LimitRanges, default NetworkPolicies, labels. Sharing: oc policy add-role-to-user view alice -n shop. Deleting one is kubectl delete ns - no undo.
Also asked: How would you design project onboarding for a multi-team OpenShift cluster? · How do you give a colleague read access to your project? · Why does every OpenShift project get its own UID range?
Learn it: 30.4 Projects vs namespaces
What are SecurityContextConstraints in OpenShift, and what does restricted-v2 do to a pod? Mid
An SCC is a cluster-scoped object listing what a pod may ask for (root, host network, volume types, capabilities) and what it gets when it does not ask. The SCC admission plugin collects the SCCs the pod may use, tries them in order, fills in the empty fields and then validates, admits with the first that fits, and records it in the pod's openshift.io/scc annotation. Pod Security Admission only judges a pod; SCC also rewrites it.
restricted-v2, the default for every authenticated user and ServiceAccount:
runAsUserfrom the project's range (MustRunAsRange) - not the image'sUSER;- primary group 0, plus an
fsGroupfrom the project's range; - the project's SELinux MCS level;
- capabilities drop ALL (only
NET_BIND_SERVICEmay be added back), no privilege escalation, RuntimeDefault seccomp; no hostPath, no host network.
So oc rsh deploy/web id prints uid=1000650000 gid=0(root). Group 0 is the hook: images whose writable directories are group-0-writable work for any UID. PSA still runs (warn/audit synced from the SCCs), but enforcement is SCC's.
Also asked: How do you make a container image run correctly under restricted-v2? · How do you find out which SCC admitted a pod and which UID it runs as? · Compare SCCs with Kubernetes Pod Security Admission.
Learn it: 30.7 SecurityContextConstraints I: what restricted-v2 does to your pod
A vendor application must run as UID 1001. How do you allow it securely on OpenShift? Mid
First check it really needs that: most images only need arbitrary-UID fixes (group-0-owned writable directories with chmod g=u, a port above 1024, emptyDirs) and then run fine under restricted-v2.
If the fixed UID is real, the narrowest grant:
- A dedicated ServiceAccount:
oc create sa reporting -n shop. nonroot-v2(any fixed non-root UID, everything else still restricted), granted to that SA in one namespace:oc adm policy add-scc-to-user nonroot-v2 -z reporting -n shop- with-nit is a RoleBinding; without it, a ClusterRoleBinding. (Or a Role with verbuseon the SCC, for GitOps.)- In the Deployment:
serviceAccountName: reporting,runAsUser: 1001, andopenshift.io/required-scc: nonroot-v2on the pod template, so it never silently lands on another SCC. - Verify:
oc adm policy scc-subject-review -z reporting -f ..., then the pod'sopenshift.io/sccannotation.
Remember only the ServiceAccount counts for Deployment pods (the ReplicaSet controller creates them). And not anyuid on default: every pod there could run as UID 0, and with priority 10 it steals pods restricted-v2 would have admitted.
Also asked: How does OpenShift decide which SCC applies to a pod? · Why is granting anyuid to the default ServiceAccount a bad fix? · How do you read an "unable to validate against any security context constraint" event?
Learn it: 30.9 SecurityContextConstraints II: selection, rejections, and fixing images
Users get the OpenShift "Application is not available" page. Walk through your troubleshooting. Mid
That page is an HTTP 503 from the router (HAProxy) - your pods never saw the request. The router sends traffic directly to pod IPs from the Service's endpoints, so the page lists its own causes, and I check them in order:
- The host does not exist -
oc get route: is there an admitted route with that host? Typo, wrong project, orHostAlreadyClaimed(an older route elsewhere owns it). - No matching path - a route with
path: /apidoes not serve/. - No usable endpoint:
oc get endpoints- no ready pods, a Service selector matching nothing, or the classic: the Route'stargetPortis the pod port (or the Service port's name), not the Service port.--port=80against endpoints on 8080 = 503;oc describe routeshows "Endpoint Port: 80". - TLS variants:
http://to a TLS route with insecure policyNone, or a reencrypt route whose backend does not speak TLS.
Different codes, different places: a router 504 = the backend was slower than the router timeout (30 s; haproxy.router.openshift.io/timeout), 502 = the backend failed, 404/500 = your app answered.
Also asked: How do you expose an application outside the cluster on OpenShift? · What is the difference between a Route and an Ingress? · How would you use Routes for a blue-green or canary release?
Learn it: 30.13 Routes and the router
Explain the three TLS termination types for OpenShift Routes and when you use each. Mid
The question is where TLS ends:
- edge - client -TLS-> router -HTTP-> pod. The router holds the certificate (its default
*.appswildcard, or one on the Route). The default for web apps and APIs.insecureEdgeTerminationPolicy:None(http gets a 503),Redirect(302 to https),Allow. - passthrough - the router only reads SNI and forwards encrypted bytes; the pod presents its own certificate. For mTLS or when the router must never see plaintext. No paths, cookies or HTTP timeouts; the backend must speak TLS (else
wrong version number). - reencrypt - client -TLS-> router -TLS-> pod. Router features and encryption on the pod network. The router verifies the pod's certificate - easiest with the service CA: annotate the Service with
service.beta.openshift.io/serving-cert-secret-name, mount the Secret, and the router trusts it with no flags.
Common mismatches: edge to an HTTPS pod gives Client sent an HTTP request to an HTTPS server; reencrypt to an HTTP or untrusted pod gives a 503. And trust the cluster's ingress CA with oc extract + curl --cacert, not -k.
Also asked: How would you implement end-to-end encryption for a service on OpenShift with minimal effort? · What does the service CA do? · Why does curl get a 503 on http:// while the browser works on https://?
What are BuildConfigs and ImageStreams in OpenShift? Mid
- A BuildConfig (
bc) is the recipe for building an image on the cluster: source (Git URL, ref, context dir), strategy (Source = S2I, Docker = a Dockerfile), output (an ImageStreamTag) and triggers (config change, builder-image change, webhooks). Each run is a numbered Build in a build pod;oc logs -f bc/NAMEfollows it; a failed one carries a reason (FetchSourceFailed,GenericBuildFailed,PushImageToRegistryFailed). - S2I builds without a Dockerfile: a language builder image runs its
assemblescript on your source. Its images are UBI-based and arbitrary-UID ready.oc new-app nodejs~<repo>creates ImageStream + BuildConfig + Deployment + Service. - An ImageStream (
is) is a named set of tags, each pointing at an image by digest, with history. Workloads following a tag always get an immutable reference; rollback = re-point a tag (oc tag);oc import-imagetracks external images. An image trigger rolls a Deployment when its tag moves.
Images live in the internal registry (image-registry.openshift-image-registry.svc:5000/<project>/<stream>); pulls across projects need system:image-puller. Many teams build in CI instead and only deploy on OpenShift.
Also asked: Compare building images with S2I on OpenShift against building in an external CI pipeline. · A build fails with FetchSourceFailed. What do you check? · How can a pod in one project pull an image from another project's ImageStream?
What is the difference between a DeploymentConfig and a Deployment, and how do you migrate without downtime? Mid
A DeploymentConfig (apps.openshift.io/v1) is OpenShift's older, deprecated since 4.14 version: it manages ReplicationControllers (<name>-<revision>) through a deployer pod, has built-in ConfigChange and ImageChange triggers and pre/mid/post lifecycle hooks. A Deployment (apps/v1) manages ReplicaSets from the HA controller manager, with matchLabels selectors, pause/resume and progressDeadlineSeconds. The deployer pod is the weakness: a rollout can stall when it cannot be scheduled.
Mapping: Rolling -> RollingUpdate; the ImageChange trigger -> oc set triggers deploy/NAME --from-image=IS:TAG -c CONTAINER (an annotation); lifecycle hooks -> a Job or init container; ConfigChange -> nothing (Deployments always roll on template change).
Migration that never empties the Service:
- Create the Deployment next to the DC with the same pod template and labels; make sure the Service selects a label both share (not
deploymentconfig=...). - Wait until it is Ready and in the endpoints.
oc scale dc/NAME --replicas=0, thenoc delete dc NAME.- Move HPAs, PDBs, Routes and monitoring that used the DC's labels.
Also asked: How would you plan the migration of many DeploymentConfigs across teams? · How do you give a Deployment the image-change trigger a DeploymentConfig had? · What replaces DeploymentConfig lifecycle hooks on a Deployment?
Learn it: 30.21 DeploymentConfig vs Deployment
An operator was installed but the CRD fields its documentation mentions do not exist. What is going on? Mid
Almost always the operator is not at the version you think - an upgrade is waiting for approval in OLM (the Operator Lifecycle Manager).
The four objects in the operator's namespace:
- OperatorGroup - which namespaces it watches (exactly one per namespace).
- Subscription - package, catalog, channel, and approval Automatic or Manual.
- InstallPlan - OLM's plan for one version (CSV + CRDs + RBAC).
- ClusterServiceVersion (CSV) - the installed version and its phase.
Check: oc get sub,ip,csv -n <ns>. An InstallPlan with APPROVED false and the Subscription condition InstallPlanPending / RequiresApproval means a Manual upgrade (or install) never happened - the CRDs are still the old version's. Approve the plan in status.installPlanRef (oc patch installplan ... --type merge -p '{"spec":{"approved":true}}'), in a change window, after testing on non-prod. Older unapproved plans are just leftovers.
Other causes: the channel does not include that version yet (change channel after reading upgrade notes), or the CSV is Failed (NoOperatorGroup, TooManyOperatorGroups, UnsupportedOperatorGroup). Manual approval is the norm in banks - and the price is that a waiting plan blocks the channel.
Also asked: What is the Operator Lifecycle Manager? · Why would a bank set operator subscriptions to Manual approval? · How do you uninstall an operator installed through OLM?
Learn it: 30.23 Operators, OperatorHub and OLM
How are OpenShift clusters upgraded, and how do you run one safely? Mid
You ask for a version; the cluster does the rest. A release image lists every platform component's image; the Cluster Version Operator applies it operator by operator; last, the Machine Config Operator updates RHCOS and the kubelet, draining and rebooting nodes one at a time per MachineConfigPool. No manual etcd, apiserver or kernel steps.
Safely:
oc adm upgrade- current version, channel, recommended updates from Red Hat's update graph (conditional ones come with a risk note). Production uses stable-4.x; even minors are EUS.- Before a minor: switch channel (
oc adm upgrade channel stable-4.22), read release notes for removed APIs (an admin-ack may be required), check every add-on operator supports the target, test on non-prod first. - Check health:
oc get coall Available, none Degraded; PDBs that allow drains. oc adm upgrade --to=<version>(or--to-latest), then followoc get clusterversion,oc get co(a stuck or Degraded operator -oc describe co),oc get mcpand nodes goingSchedulingDisabled.
Upgrades are forward only - no downgrade; the safety nets are non-prod testing and etcd backups. If a node misbehaves: oc debug node/x + chroot /host, and oc adm must-gather for support.
Also asked: A worker node is NotReady on OpenShift. How do you investigate without SSH? · What are OpenShift update channels and which would you use in production? · What is must-gather and when do you use it?
Learn it: 30.26 Cluster operations: upgrades, cluster operators, node debugging, must-gather
What are the most useful oc commands for day-to-day work on OpenShift, coming from kubectl? Mid
oc contains all of kubectl - get, describe, logs, apply, rollout, auth can-i work unchanged. What changes:
- Access:
oc login https://api.<cluster>:6443,oc whoami(-t,--show-console). - Namespaces:
oc new-project x(self-service, you become admin),oc project x,oc projects(only yours). - Deploy:
oc new-app --image=...oroc new-app builder~repo(S2I build + ImageStream + Deployment + Service). - Expose:
oc expose svc/x(a Route),oc create route edge|passthrough|reencrypt. - Debug:
oc rsh pod(a shell),oc debug node/x+chroot /host,oc status(the Route -> Service -> Deployment -> Build chain and what is broken). - Access for others:
oc policy add-role-to-user edit alice. - Security: the pod's
openshift.io/sccannotation andoc adm policy scc-subject-review. - Cluster:
oc get co,oc adm upgrade,oc adm must-gather.
And a trap: oc get all is a category - it leaves out ConfigMaps, Secrets, PVCs, RoleBindings, NetworkPolicies and ServiceAccounts. List those explicitly.
Also asked: You join a team that runs everything on OpenShift, coming from AKS. What do you check in your first week? · What does oc status show that kubectl get all does not? · What are OpenShift Templates and what replaced them?
Learn it: 30.29 A working day on OpenShift: the translation table
Practise these answers with flashcards and labs Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.