Kubernetes: Scheduling, Health & Security: interview questions
The question you are most likely to get for each topic, a model answer, and what else comes up. From chapter 17 of the course.
A pod keeps restarting. How do you diagnose it? Mid
k describe pod POD first: Last State, its exit code and reason, and the events. Then sort by cause:
OOMKilled, 137 - the container went overlimits.memory: compare usage (kubectl top podover time) with the limit; for a JVM, the heap setting against the limit.Errorwith 143 or 137, plusLiveness probe failedevents - the liveness probe kills it: probe too strict, too early (a slow starter needs a startup probe), or checking a dependency instead of the process.- Exit 1 or similar, no probe events - the app itself exits:
k logs POD --previousfor the crashed run. - Evicted pods instead of restarts - node memory pressure: the QoS class decides who goes first.
And look around it: CPU throttling (cpu.stat nr_throttled) slows startup into probe timeouts, and CrashLoopBackOff is only the kubelet waiting between restarts, never the cause.
Also asked: What is the difference between resource requests and limits? · What is the difference between liveness, readiness and startup probes? · What is the difference between a Role and a ClusterRole?
Explain resource requests and limits in Kubernetes. Junior
Two numbers per resource, read by different components:
- Requests - what the scheduler reserves. It adds up the requests of the pods on a node and compares them with the node's allocatable; it never looks at actual usage. On the node,
requests.cpubecomescpu.weight: a fair share when the CPU is contended. Memory requests also rank pods for eviction. - Limits - hard ceilings the kernel enforces through cgroups:
limits.cpubecomescpu.max(a quota per 100 ms period - over it you are throttled),limits.memorybecomesmemory.max(over it the container is OOM-killed, exit 137).
Consequences: Insufficient cpu on an idle cluster means requested, not used; a pod without requests is invisible to the scheduler and gets overpacked and evicted first; limits are not reserved, so they may add up past 100% (overcommit). Units: 250m = a quarter core, 256Mi = mebibytes - 128m memory is 0.128 bytes. kubectl describe node shows the ledger.
Also asked: The cluster is only 20% utilised, yet new pods do not schedule. Why? · How do you choose request and limit values for a service? · What happens if you set only a limit and no request?
Learn it: 17.1 Requests and limits: what each one actually controls
What are the Kubernetes QoS classes, and which pods get evicted first? Junior
The QoS class is computed from every container's resources (never declared) and written to status.qosClass:
- Guaranteed - every container has CPU and memory limits, with requests equal to them.
- Burstable - at least one request or limit, but not Guaranteed.
- BestEffort - no requests or limits at all.
When a node runs low (below the eviction threshold, e.g. memory.available<100Mi), the kubelet evicts pods, ranked by: is usage above the requests, then priority, then how far above. So BestEffort (request 0, always "above") goes first, Burstable over its requests next, Guaranteed last. The kubelet also sets oom_score_adj by class (Guaranteed -997, BestEffort 1000) so the kernel agrees if it acts first.
Keep that apart from a container hitting its own memory limit: a cgroup OOM kill, the container restarts with exit 137, whatever the class.
Also asked: How do you make a critical service less likely to be evicted? · What is the difference between an eviction and an OOM kill? · Why can one sidecar change a pod's QoS class?
Learn it: 17.3 QoS classes, oom_score_adj, and who gets evicted first
What is the difference between hitting a CPU limit and hitting a memory limit? Mid
CPU is compressible, memory is not.
- Over the CPU limit - the cgroup is throttled: paused until the next 100 ms period. No restart; the symptom is latency. A service with 8 busy threads and
limits.cpu: 1uses its quota in 12.5 ms and sleeps 87.5 ms - average CPU looks like 40% while p99 spikes. Prove it fromcpu.stat:nr_throttled / nr_periods, two samples apart. Fix: raise or remove the CPU limit (keep the request), or make the code cheaper. - Over the memory limit - the kernel OOM-kills a process in the container: restart, exit 137,
Reason: OOMKilledin Last State. Fix: raise the limit or use less.
For a JVM, the classic 137 is a heap allowed to grow to the limit with no room for metaspace, thread stacks and the rest: size the limit as heap + non-heap + margin, and express the heap as a percentage (-XX:MaxRAMPercentage=75) so they cannot disagree.
Also asked: A Java service restarts every few hours with exit code 137. How do you diagnose it? · Should you set CPU limits on latency-sensitive services? · Why does kubectl top not show CPU throttling?
Learn it: 17.6 CPU throttling vs memory OOMKill: two failure modes, two fixes
What is the difference between a LimitRange and a ResourceQuota? Junior
Both are namespaced admission rules, with different jobs:
- LimitRange - per container (or pod, PVC): fills in default requests and limits when they are missing, and enforces min/max. Applied when the pod is created.
- ResourceQuota - a total budget for the namespace: the sum of
requests.cpu,limits.memory, the number of pods, PVCs, Services...kubectl get quotashows Used against Hard.
They come as a pair: once a quota covers a resource, every new pod must declare it ("must specify limits.cpu"), and the LimitRange's defaults are what fill it in. Going over gives "exceeded quota".
The trap when a Deployment creates the pods: you see no error - the ReplicaSet controller gets it. A Deployment stuck below its replica count with no Pending pods: read the ReplicaSet's events. kubectl describe namespace shows the quotas and limit ranges in one place.
Also asked: A team says they cannot deploy anymore. How do you investigate? · Why must a namespace with a ResourceQuota usually also have a LimitRange? · Where do you see the error when a quota blocks a Deployment's pods?
Learn it: 17.9 LimitRange and ResourceQuota: per-pod defaults vs namespace totals
How do you make a pod run only on specific nodes? Junior
Label the nodes, then select them:
kubectl label node worker-1 disktype=ssdnodeSelectorin the pod spec:disktype: ssd- every listed label must match. The simple, hard rule.- Node affinity for more:
requiredDuringSchedulingIgnoredDuringExecutionwithmatchExpressions(In,NotIn,Exists,Gt...) - terms are ORed, expressions inside a term ANDed;preferred...with aweightfor "prefer, but run anyway".
Nothing matching leaves the pod Pending: "didn't match Pod's node affinity/selector".
Two things to say: IgnoredDuringExecution means the rule is checked only at scheduling - remove a label later and running pods stay, until the next restart cannot be placed. And affinity only attracts; it does not keep other pods off those nodes - that needs a taint.
Also asked: What does IgnoredDuringExecution mean, and what risk does it create? · How would you run a workload only on arm64 nodes in a mixed cluster? · What does setting spec.nodeName do?
Learn it: 17.11 nodeSelector and node affinity
What are taints and tolerations? Junior
A taint is on the node and repels pods: kubectl taint node worker-2 dedicated=batch:NoSchedule (key, value, effect; a trailing - removes it). A toleration is on the pod and permits it on a node with a matching taint.
The effects:
NoSchedule- new pods that do not tolerate it are not placed there; running pods stay.PreferNoSchedule- the scheduler tries to avoid the node.NoExecute- also evicts running pods that do not tolerate it (optionally aftertolerationSeconds).
Examples you already have: the control plane's node-role.kubernetes.io/control-plane:NoSchedule, and the built-in not-ready/unreachable NoExecute taints that every pod tolerates for 300 s - the five minutes before pods leave a dead node.
The key point: a toleration is permission, not attraction. A dedicated node needs the taint (keeps others off), the toleration (lets yours on) and a nodeSelector or affinity (keeps yours there).
Also asked: How do you dedicate a set of nodes to one workload? · A node dies. What happens to its pods, and when? · What does kubectl cordon do to a node?
Learn it: 17.13 Taints and tolerations: the node repels, the pod tolerates
How do you make sure replicas of a Deployment do not all end up on the same node or zone? Mid
The default spreading is only a soft score, so state it:
- Pod anti-affinity with
topologyKey: kubernetes.io/hostname(ortopology.kubernetes.io/zone): "never put me where a pod withapp=webalready runs". Required = at most one per domain - replicas beyond the number of nodes stay Pending, and it can block a rolling update's surge pod. - Topology spread constraints - built for this:
maxSkew: 1,topologyKey,whenUnsatisfiable: DoNotSchedule(hard) orScheduleAnyway(soft),labelSelectorfor your own pods. Arithmetic, not binary: 3 replicas on 2 nodes become 2 + 1, so any count runs.
A common production default: hard across zones, soft across nodes. Watch the kubeadm trap: with the default nodeTaintsPolicy: Ignore, the tainted control-plane node counts as an empty domain and blocks the spread - set nodeTaintsPolicy: Honor. And nothing rebalances running pods afterwards.
Also asked: What is the difference between pod affinity and pod anti-affinity? · What is maxSkew in a topology spread constraint? · How would you design pod placement for a service that must survive a zone outage?
Learn it: 17.15 Pod affinity, anti-affinity and topology spread
A pod is Pending. How do you find out why? Junior
k describe pod POD and read the FailedScheduling event as a table:
0/3 nodes are available: 1 node(s) had untolerated taint {...control-plane: },
1 Insufficient cpu, 1 node(s) didn't match Pod's node affinity/selector.
0/3- three nodes considered, none passed.- One reason per node - the first filter that rejected it: a taint, not enough requested CPU or memory against allocatable, labels/affinity, anti-affinity or spread rules, an unbound volume.
preemption:- whether evicting lower-priority pods could help ("not helpful" for taints and labels).
Fix the reason for each node - and if the pod is still Pending, read the new message: the next filter in line may now be the one failing. No event at all? Check the pod has no nodeName and the scheduler is running. With a full cluster, a PriorityClass lets important pods be scheduled first and preempt lower-priority ones.
Also asked: What is a PriorityClass, and what is preemption? · You fixed the reason in FailedScheduling but the pod is still Pending. Why? · What does "Insufficient cpu" mean on a cluster with idle CPUs?
Learn it: 17.18 Reading FailedScheduling, and PriorityClass with preemption
Explain the difference between liveness, readiness and startup probes. Junior
The kubelet runs all three against the container; each answers one question, and failing it has one consequence:
- Liveness - "is this process wedged beyond repair?" Failure: the kubelet kills and restarts the container.
- Readiness - "should this pod get traffic right now?" Failure: the pod is removed from the Service endpoints (
0/1READY). No restart, ever - and if every replica fails readiness, the Service has no endpoints. - Startup - "has the app finished starting?" Until it succeeds, liveness and readiness do not run; if it never does, the container is restarted.
Handlers: httpGet (200-399), tcpSocket, exec, grpc. Defaults: period 10 s, timeout 1 s, failureThreshold 3 - time to act is about initialDelaySeconds + failureThreshold x periodSeconds.
The Unhealthy events say why: connection refused (not listening yet), statuscode: 503 (the app said no), context deadline exceeded (slower than the timeout). Readiness also gates rollouts: a new pod that never gets Ready stops a bad release.
Also asked: Pods are Running but 0/1 Ready and the Service returns errors. What is happening? · Which probe handler would you use for a database, a gRPC service and an HTTP API? · What exit code do you see after a liveness probe kills a container?
Learn it: 17.20 Liveness, readiness, startup: what each answers and what failure does
How would you design the probes for a Java service that takes 90 seconds to start? Mid
Three probes, three jobs, each number justified:
- startupProbe -
failureThreshold x periodSecondswell above the worst startup, e.g. 30 x 5 s = 150 s, on a local health URL. Until it passes, liveness and readiness do not run, so a slow boot is never killed - and a pod that is actually up goes Ready soon after. - readinessProbe - fast (period 5 s), may include the dependencies the pod truly needs to serve: being wrong only drops it from rotation briefly.
- livenessProbe - slow (30 s+ to act) and local only: "is this process stuck", never the database. A false positive restarts a healthy JVM.
timeoutSeconds: 2- a garbage-collection pause must not count as dead.
The reason: a liveness probe that checks the database turns a slow database into a restart storm - every pod killed at once, each restart a 90-second boot. Also check the CPU limit: startup is CPU-bound, and throttling can double it.
Also asked: What is a restart storm, and how do you prevent one? · Why must a liveness probe never check a dependency? · How would you review the probes in a teammate's Deployment YAML?
Learn it: 17.22 Probe design: the restart storm, and why JVMs need startup probes
How does a Deployment rolling update work, and how do you make it safe? Junior
A template change creates a new ReplicaSet; the Deployment scales it up and the old one down within two limits:
- maxSurge - pods allowed above
replicas(rounded up) - maxUnavailable - pods allowed below
replicas(rounded down)
maxSurge: 1, maxUnavailable: 0 is zero downtime: one new pod, wait until it is available (Ready for minReadySeconds), remove one old, repeat.
Readiness is what makes it safe: a new version whose readiness never passes stalls the rollout with the old pods still serving. Without a readiness probe, a broken release replaces every pod within a minute. After progressDeadlineSeconds the Deployment reports ProgressDeadlineExceeded - but does not roll back on its own.
The verbs: kubectl rollout status --timeout (non-zero exit on failure - deploy scripts use it), history (with a kubernetes.io/change-cause annotation), undo --to-revision=N, restart, pause/resume. Undo only restores the pod template, not a ConfigMap you changed - and if git re-applies, revert there too.
Also asked: A release went out and errors spiked. What do you do in the first five minutes? · When would you use the Recreate strategy? · What does kubectl rollout undo restore, and what does it not?
Learn it: 17.24 Rolling updates: maxSurge, maxUnavailable, and the rollout verbs
What is a PodDisruptionBudget? Junior
A PDB limits how many pods of a set may be down because of a voluntary disruption - a kubectl drain, a node upgrade, the cluster autoscaler removing a node. It cannot help against crashes, node failures or kubelet evictions.
apiVersion: policy/v1
kind: PodDisruptionBudget
spec:
maxUnavailable: 1
selector: {matchLabels: {app: web}}
It works through the Eviction API: kubectl drain cordons the node and asks to evict each pod; the API server refuses when that would break the budget, and the drain retries. kubectl delete pod ignores PDBs entirely. kubectl get pdb shows ALLOWED DISRUPTIONS.
The classic self-inflicted wound: one replica with minAvailable: 1 - allowed disruptions is 0 forever, and every drain or node upgrade hangs. Run at least two replicas, prefer maxUnavailable, and consider unhealthyPodEvictionPolicy: AlwaysAllow so broken pods do not block drains.
Also asked: A node pool upgrade is stuck. How do you find and fix the cause? · What flags does kubectl drain usually need, and why? · What is the difference between a voluntary and an involuntary disruption?
Learn it: 17.26 PodDisruptionBudgets and why they matter during a drain
How does the Horizontal Pod Autoscaler work? Junior
A control loop in kube-controller-manager, every 15 seconds: read the metric of the target's pods (CPU and memory from metrics-server), compute
desiredReplicas = ceil(currentReplicas x currentValue / targetValue)
and write it into the Deployment's replica count through the scale subresource. 3 pods at 90% with a 50% target -> ceil(5.4) = 6. Within 10% of the target it does nothing.
The traps:
- Utilization is a percentage of the requests - no CPU request (on every container) means
<unknown>and no scaling. - Do not set
replicasin the manifest you apply, and do notkubectl scale- you fight the HPA. - New replicas need capacity: Pending pods if the nodes are full.
- Scale-down waits 5 minutes (
behaviorstabilization) by design.
kubectl autoscale deployment web --cpu=50% --min=2 --max=10 creates one; describe hpa conditions (AbleToScale, ScalingActive, ScalingLimited) explain what it is doing.
Also asked: An HPA is not scaling the way the team expects. How do you debug it? · Why does an HPA on CPU need resource requests? · When is CPU a bad metric for autoscaling?
Learn it: 17.28 The HorizontalPodAutoscaler
What is the difference between a Role and a ClusterRole, and between a RoleBinding and a ClusterRoleBinding? Junior
RBAC answers "may this identity do this verb on this resource (in this API group) in this namespace?". Roles say what; bindings say who and where.
- Role - rules in one namespace.
- ClusterRole - the same rules without a namespace: needed for cluster-scoped resources (nodes, PVs, namespaces) and for definitions reused in many namespaces (
view,edit,admin,cluster-admin). - RoleBinding - grants a Role or a ClusterRole to subjects (users, groups, ServiceAccounts) only in its namespace. Binding the built-in
viewClusterRole with a RoleBinding indevgives view ofdevonly. - ClusterRoleBinding - grants a ClusterRole everywhere, plus cluster-scoped resources.
RBAC is additive - no deny rules; any matching rule allows. The Forbidden message is the spec of the missing rule (user, verb, resource, API group, namespace), the most common bug is the wrong apiGroups (deployments are in apps), and list on secrets is read access to all of them.
Also asked: What are the built-in user-facing ClusterRoles, and what can each do? · Why does RBAC have no deny rules, and what does that mean for removing access? · A ServiceAccount gets "Forbidden: cannot list resource deployments in API group apps". How do you fix it properly?
Learn it: 17.30 RBAC in full: Role, ClusterRole, RoleBinding, ClusterRoleBinding
What is a ServiceAccount, and how does a pod use it? Junior
A ServiceAccount is the identity a pod uses when it calls the Kubernetes API (a controller, a CI runner, an app using a Kubernetes client library). RBAC sees it as system:serviceaccount:<namespace>:<name>.
- Every namespace has a
defaultSA, and pods that do not name one run as it. - The kubelet mounts a projected, bound, short-lived token (about an hour, rotated, tied to the pod) with the cluster CA at
/var/run/secrets/kubernetes.io/serviceaccount/; client libraries read it from there.
Good practice:
- A dedicated SA per workload that calls the API (
kubectl create serviceaccount podwatcher,serviceAccountNamein the pod template) with a Role for exactly what it needs. automountServiceAccountToken: falsefor everything that never calls the API - otherwise any permission someone grants todefaultreaches every pod.
kubectl create token SA --duration=10m issues a token by hand; 401 means the token is bad or expired, 403 means RBAC said no.
Also asked: How do you give a CI job access to deploy to a cluster? · Why set automountServiceAccountToken: false, and where? · What is inside a ServiceAccount token, and is it encrypted?
Learn it: 17.33 ServiceAccounts, tokens, and automountServiceAccountToken
How do you check whether a user or ServiceAccount can perform an action? Junior
Ask the authorizer directly with kubectl auth can-i - it prints yes or no and exits 0 or 1, so it is scriptable:
kubectl auth can-i create deployments -n prod
kubectl auth can-i list secrets -n dev --as=jane
kubectl auth can-i patch deployments -n prod --as=system:serviceaccount:ci:deployer
kubectl auth can-i get pods --subresource=log -n dev --as=jane
kubectl auth can-i --list -n dev --as=system:serviceaccount:dev:app
--as impersonates (you need the impersonate verb, which cluster-admin has), so you do not need the other identity's credentials. --list shows everything that identity may do in the namespace.
From a Forbidden message to the fix: write the rule in the same words - cannot patch resource "deployments" in API group "apps" in the namespace "prod" becomes apiGroups: ["apps"], resources: ["deployments"], verbs: ["patch"] in a Role in prod, bound to that ServiceAccount. Then can-i again. A 401 is never fixed with a Role.
Also asked: A controller in the cluster logs Forbidden errors. How do you fix it? · How would you find out who can delete pods in a production namespace? · What is the difference between a 401 and a 403 from the API server?
Which securityContext settings would you require for every application container, and why? Mid
Each closes a real hole if someone gets code running in the container:
runAsNonRoot: true(+ a numericrunAsUser) - most images run as root, and root in a container is root on the node's kernel. The kubelet refuses to start a root container (CreateContainerConfigError).allowPrivilegeEscalation: false- the no_new_privs flag: setuid binaries cannot raise privileges again.capabilities: {drop: ["ALL"]}- containers get 14 kernel capabilities by default; most apps need none.readOnlyRootFilesystem: true- nobody can drop a binary or edit config; give the app anemptyDirwhere it must write (/tmp, caches).seccompProfile: {type: RuntimeDefault}- filters dangerous system calls; Kubernetes does not apply it by default.
And never privileged: true for an app. Verify from /proc/1/status (Uid, CapEff, NoNewPrivs, Seccomp). The failures tell you what to fix: Permission denied (wrong uid) vs Read-only file system (needs a volume). That block is exactly the restricted Pod Security Standard.
Also asked: A container fails after you enforced readOnlyRootFilesystem and runAsNonRoot. How do you help the team? · What is a Linux capability, and why drop them all? · What does allowPrivilegeEscalation: false do?
Learn it: 17.37 securityContext: the fields worth setting every time
What are the Pod Security Standards, and how would you roll out restricted without breaking teams? Mid
Pod Security Admission is built into the API server and checks pods against three levels, set per namespace with labels:
- privileged - no restrictions (the default).
- baseline - blocks known escalations: privileged, host namespaces, hostPath, extra capabilities.
- restricted - baseline plus non-root,
allowPrivilegeEscalation: false, drop ALL capabilities, a seccomp profile, safe volume types only.
Three modes: enforce rejects, warn tells the client, audit logs. Roll out in steps:
enforce: baseline+warn: restricted(andaudit: restricted) - nothing breaks, everyone sees what would.- Preview:
kubectl label ns shop pod-security.kubernetes.io/enforce=restricted --dry-run=serverlists the violating pods. - Fix the manifests, then enforce restricted; pin
enforce-versionso upgrades do not change the rules.
The trap: enforce checks pods, so a Deployment is accepted and its ReplicaSet then cannot create pods - which is why warn (checked on workloads) matters. Running pods are never evicted, only their replacements rejected.
Also asked: What are the limits of Pod Security Admission, and what would you add on top? · Why does a Deployment get accepted in a restricted namespace and still create no pods? · What is the difference between the enforce, warn and audit modes?
Learn it: 17.39 Pod Security Standards and Pod Security Admission
How would you manage application secrets on a Kubernetes platform in a regulated environment? Mid
First, who can read a Secret today - four layers:
- RBAC -
get,listandwatchon secrets (list returns full objects), and anyone who can create pods in the namespace can mount any secret there. - etcd - without encryption at rest, secrets are stored as plain base64, in every backup.
- The node - root reads mounted secrets from tmpfs.
- The process - env vars leak through
/proc/PID/environ, crash dumps and debug endpoints.
Then the controls:
- Encryption at rest with a KMS provider (
EncryptionConfiguration), and rewrite existing secrets so they are encrypted; verify in etcd (k8s:enc:kms:v2:). - The source of truth outside the cluster: a secrets manager (a vault) with its own policies, audit and rotation, delivered by the Secrets Store CSI driver using the pod's identity, or synced by an operator.
- Tight RBAC: no human
list secretsin production; audit-log secret reads. - Mounted files over env vars; never Secret manifests in git unless encrypted (SOPS, Sealed Secrets);
immutable: truewith versioned names for rotation.
Also asked: Are Kubernetes Secrets secure? · How do you enable encryption at rest for Secrets, and verify it works? · Why is permission to create pods also access to secrets?
Learn it: 17.41 Secrets, and why base64 is not a security measure
How do you give a new person access to a kubeadm cluster with a client certificate? Mid
Kubernetes has no User objects: the certificate is the user. The API server trusts certificates signed by the cluster CA; CN becomes the username, each O a group.
- Key and request:
openssl genrsa -out jane.key 2048,openssl req -new -key jane.key -out jane.csr -subj "/CN=jane/O=dev". - Submit a CertificateSigningRequest object:
request=base64 -w0 jane.csr,signerName: kubernetes.io/kube-apiserver-client,usages: [client auth], a shortexpirationSeconds. - Check the subject (
kubectl describe csr jane), thenkubectl certificate approve jane; the controller-manager issues it. - Extract:
kubectl get csr jane -o jsonpath='{.status.certificate}' | base64 -d > jane.crt. - kubeconfig:
kubectl config set-credentials jane --client-key=jane.key --client-certificate=jane.crt --embed-certs=true, plusset-context. - Authorize: a RoleBinding, preferably to the group (
--group=dev). - Test with
kubectl --context jane auth whoamiandget pods: 401 = certificate problem, 403 = RBAC.
Certificates cannot be revoked, so keep them short-lived - and never O=system:masters.
Also asked: A user's kubectl returns errors. How do you tell whether it is the certificate or RBAC? · Why are client certificates risky for human access in production? · What does a CSR stuck in Approved without Issued tell you?
Learn it: 17.43 Users are certificates: the CSR flow end to end
Practise these answers with flashcards and labs Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.