OnCallReady

Kubernetes: Scheduling, Health & Security: interview questions

The question you are most likely to get for each topic, a model answer, and what else comes up. From chapter 17 of the course.

A pod keeps restarting. How do you diagnose it? Mid

k describe pod POD first: Last State, its exit code and reason, and the events. Then sort by cause:

And look around it: CPU throttling (cpu.stat nr_throttled) slows startup into probe timeouts, and CrashLoopBackOff is only the kubelet waiting between restarts, never the cause.

Also asked: What is the difference between resource requests and limits? · What is the difference between liveness, readiness and startup probes? · What is the difference between a Role and a ClusterRole?

Explain resource requests and limits in Kubernetes. Junior

Two numbers per resource, read by different components:

Consequences: Insufficient cpu on an idle cluster means requested, not used; a pod without requests is invisible to the scheduler and gets overpacked and evicted first; limits are not reserved, so they may add up past 100% (overcommit). Units: 250m = a quarter core, 256Mi = mebibytes - 128m memory is 0.128 bytes. kubectl describe node shows the ledger.

Also asked: The cluster is only 20% utilised, yet new pods do not schedule. Why? · How do you choose request and limit values for a service? · What happens if you set only a limit and no request?

Learn it: 17.1 Requests and limits: what each one actually controls

What are the Kubernetes QoS classes, and which pods get evicted first? Junior

The QoS class is computed from every container's resources (never declared) and written to status.qosClass:

When a node runs low (below the eviction threshold, e.g. memory.available<100Mi), the kubelet evicts pods, ranked by: is usage above the requests, then priority, then how far above. So BestEffort (request 0, always "above") goes first, Burstable over its requests next, Guaranteed last. The kubelet also sets oom_score_adj by class (Guaranteed -997, BestEffort 1000) so the kernel agrees if it acts first.

Keep that apart from a container hitting its own memory limit: a cgroup OOM kill, the container restarts with exit 137, whatever the class.

Also asked: How do you make a critical service less likely to be evicted? · What is the difference between an eviction and an OOM kill? · Why can one sidecar change a pod's QoS class?

Learn it: 17.3 QoS classes, oom_score_adj, and who gets evicted first

What is the difference between hitting a CPU limit and hitting a memory limit? Mid

CPU is compressible, memory is not.

For a JVM, the classic 137 is a heap allowed to grow to the limit with no room for metaspace, thread stacks and the rest: size the limit as heap + non-heap + margin, and express the heap as a percentage (-XX:MaxRAMPercentage=75) so they cannot disagree.

Also asked: A Java service restarts every few hours with exit code 137. How do you diagnose it? · Should you set CPU limits on latency-sensitive services? · Why does kubectl top not show CPU throttling?

Learn it: 17.6 CPU throttling vs memory OOMKill: two failure modes, two fixes

What is the difference between a LimitRange and a ResourceQuota? Junior

Both are namespaced admission rules, with different jobs:

They come as a pair: once a quota covers a resource, every new pod must declare it ("must specify limits.cpu"), and the LimitRange's defaults are what fill it in. Going over gives "exceeded quota".

The trap when a Deployment creates the pods: you see no error - the ReplicaSet controller gets it. A Deployment stuck below its replica count with no Pending pods: read the ReplicaSet's events. kubectl describe namespace shows the quotas and limit ranges in one place.

Also asked: A team says they cannot deploy anymore. How do you investigate? · Why must a namespace with a ResourceQuota usually also have a LimitRange? · Where do you see the error when a quota blocks a Deployment's pods?

Learn it: 17.9 LimitRange and ResourceQuota: per-pod defaults vs namespace totals

How do you make a pod run only on specific nodes? Junior

Label the nodes, then select them:

Nothing matching leaves the pod Pending: "didn't match Pod's node affinity/selector".

Two things to say: IgnoredDuringExecution means the rule is checked only at scheduling - remove a label later and running pods stay, until the next restart cannot be placed. And affinity only attracts; it does not keep other pods off those nodes - that needs a taint.

Also asked: What does IgnoredDuringExecution mean, and what risk does it create? · How would you run a workload only on arm64 nodes in a mixed cluster? · What does setting spec.nodeName do?

Learn it: 17.11 nodeSelector and node affinity

What are taints and tolerations? Junior

A taint is on the node and repels pods: kubectl taint node worker-2 dedicated=batch:NoSchedule (key, value, effect; a trailing - removes it). A toleration is on the pod and permits it on a node with a matching taint.

The effects:

Examples you already have: the control plane's node-role.kubernetes.io/control-plane:NoSchedule, and the built-in not-ready/unreachable NoExecute taints that every pod tolerates for 300 s - the five minutes before pods leave a dead node.

The key point: a toleration is permission, not attraction. A dedicated node needs the taint (keeps others off), the toleration (lets yours on) and a nodeSelector or affinity (keeps yours there).

Also asked: How do you dedicate a set of nodes to one workload? · A node dies. What happens to its pods, and when? · What does kubectl cordon do to a node?

Learn it: 17.13 Taints and tolerations: the node repels, the pod tolerates

How do you make sure replicas of a Deployment do not all end up on the same node or zone? Mid

The default spreading is only a soft score, so state it:

A common production default: hard across zones, soft across nodes. Watch the kubeadm trap: with the default nodeTaintsPolicy: Ignore, the tainted control-plane node counts as an empty domain and blocks the spread - set nodeTaintsPolicy: Honor. And nothing rebalances running pods afterwards.

Also asked: What is the difference between pod affinity and pod anti-affinity? · What is maxSkew in a topology spread constraint? · How would you design pod placement for a service that must survive a zone outage?

Learn it: 17.15 Pod affinity, anti-affinity and topology spread

A pod is Pending. How do you find out why? Junior

k describe pod POD and read the FailedScheduling event as a table:

0/3 nodes are available: 1 node(s) had untolerated taint {...control-plane: },
1 Insufficient cpu, 1 node(s) didn't match Pod's node affinity/selector.

Fix the reason for each node - and if the pod is still Pending, read the new message: the next filter in line may now be the one failing. No event at all? Check the pod has no nodeName and the scheduler is running. With a full cluster, a PriorityClass lets important pods be scheduled first and preempt lower-priority ones.

Also asked: What is a PriorityClass, and what is preemption? · You fixed the reason in FailedScheduling but the pod is still Pending. Why? · What does "Insufficient cpu" mean on a cluster with idle CPUs?

Learn it: 17.18 Reading FailedScheduling, and PriorityClass with preemption

Explain the difference between liveness, readiness and startup probes. Junior

The kubelet runs all three against the container; each answers one question, and failing it has one consequence:

Handlers: httpGet (200-399), tcpSocket, exec, grpc. Defaults: period 10 s, timeout 1 s, failureThreshold 3 - time to act is about initialDelaySeconds + failureThreshold x periodSeconds.

The Unhealthy events say why: connection refused (not listening yet), statuscode: 503 (the app said no), context deadline exceeded (slower than the timeout). Readiness also gates rollouts: a new pod that never gets Ready stops a bad release.

Also asked: Pods are Running but 0/1 Ready and the Service returns errors. What is happening? · Which probe handler would you use for a database, a gRPC service and an HTTP API? · What exit code do you see after a liveness probe kills a container?

Learn it: 17.20 Liveness, readiness, startup: what each answers and what failure does

How would you design the probes for a Java service that takes 90 seconds to start? Mid

Three probes, three jobs, each number justified:

The reason: a liveness probe that checks the database turns a slow database into a restart storm - every pod killed at once, each restart a 90-second boot. Also check the CPU limit: startup is CPU-bound, and throttling can double it.

Also asked: What is a restart storm, and how do you prevent one? · Why must a liveness probe never check a dependency? · How would you review the probes in a teammate's Deployment YAML?

Learn it: 17.22 Probe design: the restart storm, and why JVMs need startup probes

How does a Deployment rolling update work, and how do you make it safe? Junior

A template change creates a new ReplicaSet; the Deployment scales it up and the old one down within two limits:

maxSurge: 1, maxUnavailable: 0 is zero downtime: one new pod, wait until it is available (Ready for minReadySeconds), remove one old, repeat.

Readiness is what makes it safe: a new version whose readiness never passes stalls the rollout with the old pods still serving. Without a readiness probe, a broken release replaces every pod within a minute. After progressDeadlineSeconds the Deployment reports ProgressDeadlineExceeded - but does not roll back on its own.

The verbs: kubectl rollout status --timeout (non-zero exit on failure - deploy scripts use it), history (with a kubernetes.io/change-cause annotation), undo --to-revision=N, restart, pause/resume. Undo only restores the pod template, not a ConfigMap you changed - and if git re-applies, revert there too.

Also asked: A release went out and errors spiked. What do you do in the first five minutes? · When would you use the Recreate strategy? · What does kubectl rollout undo restore, and what does it not?

Learn it: 17.24 Rolling updates: maxSurge, maxUnavailable, and the rollout verbs

What is a PodDisruptionBudget? Junior

A PDB limits how many pods of a set may be down because of a voluntary disruption - a kubectl drain, a node upgrade, the cluster autoscaler removing a node. It cannot help against crashes, node failures or kubelet evictions.

apiVersion: policy/v1
kind: PodDisruptionBudget
spec:
  maxUnavailable: 1
  selector: {matchLabels: {app: web}}

It works through the Eviction API: kubectl drain cordons the node and asks to evict each pod; the API server refuses when that would break the budget, and the drain retries. kubectl delete pod ignores PDBs entirely. kubectl get pdb shows ALLOWED DISRUPTIONS.

The classic self-inflicted wound: one replica with minAvailable: 1 - allowed disruptions is 0 forever, and every drain or node upgrade hangs. Run at least two replicas, prefer maxUnavailable, and consider unhealthyPodEvictionPolicy: AlwaysAllow so broken pods do not block drains.

Also asked: A node pool upgrade is stuck. How do you find and fix the cause? · What flags does kubectl drain usually need, and why? · What is the difference between a voluntary and an involuntary disruption?

Learn it: 17.26 PodDisruptionBudgets and why they matter during a drain

How does the Horizontal Pod Autoscaler work? Junior

A control loop in kube-controller-manager, every 15 seconds: read the metric of the target's pods (CPU and memory from metrics-server), compute

desiredReplicas = ceil(currentReplicas x currentValue / targetValue)

and write it into the Deployment's replica count through the scale subresource. 3 pods at 90% with a 50% target -> ceil(5.4) = 6. Within 10% of the target it does nothing.

The traps:

kubectl autoscale deployment web --cpu=50% --min=2 --max=10 creates one; describe hpa conditions (AbleToScale, ScalingActive, ScalingLimited) explain what it is doing.

Also asked: An HPA is not scaling the way the team expects. How do you debug it? · Why does an HPA on CPU need resource requests? · When is CPU a bad metric for autoscaling?

Learn it: 17.28 The HorizontalPodAutoscaler

What is the difference between a Role and a ClusterRole, and between a RoleBinding and a ClusterRoleBinding? Junior

RBAC answers "may this identity do this verb on this resource (in this API group) in this namespace?". Roles say what; bindings say who and where.

RBAC is additive - no deny rules; any matching rule allows. The Forbidden message is the spec of the missing rule (user, verb, resource, API group, namespace), the most common bug is the wrong apiGroups (deployments are in apps), and list on secrets is read access to all of them.

Also asked: What are the built-in user-facing ClusterRoles, and what can each do? · Why does RBAC have no deny rules, and what does that mean for removing access? · A ServiceAccount gets "Forbidden: cannot list resource deployments in API group apps". How do you fix it properly?

Learn it: 17.30 RBAC in full: Role, ClusterRole, RoleBinding, ClusterRoleBinding

What is a ServiceAccount, and how does a pod use it? Junior

A ServiceAccount is the identity a pod uses when it calls the Kubernetes API (a controller, a CI runner, an app using a Kubernetes client library). RBAC sees it as system:serviceaccount:<namespace>:<name>.

Good practice:

kubectl create token SA --duration=10m issues a token by hand; 401 means the token is bad or expired, 403 means RBAC said no.

Also asked: How do you give a CI job access to deploy to a cluster? · Why set automountServiceAccountToken: false, and where? · What is inside a ServiceAccount token, and is it encrypted?

Learn it: 17.33 ServiceAccounts, tokens, and automountServiceAccountToken

How do you check whether a user or ServiceAccount can perform an action? Junior

Ask the authorizer directly with kubectl auth can-i - it prints yes or no and exits 0 or 1, so it is scriptable:

kubectl auth can-i create deployments -n prod
kubectl auth can-i list secrets -n dev --as=jane
kubectl auth can-i patch deployments -n prod --as=system:serviceaccount:ci:deployer
kubectl auth can-i get pods --subresource=log -n dev --as=jane
kubectl auth can-i --list -n dev --as=system:serviceaccount:dev:app

--as impersonates (you need the impersonate verb, which cluster-admin has), so you do not need the other identity's credentials. --list shows everything that identity may do in the namespace.

From a Forbidden message to the fix: write the rule in the same words - cannot patch resource "deployments" in API group "apps" in the namespace "prod" becomes apiGroups: ["apps"], resources: ["deployments"], verbs: ["patch"] in a Role in prod, bound to that ServiceAccount. Then can-i again. A 401 is never fixed with a Role.

Also asked: A controller in the cluster logs Forbidden errors. How do you fix it? · How would you find out who can delete pods in a production namespace? · What is the difference between a 401 and a 403 from the API server?

Learn it: 17.35 kubectl auth can-i as a debugging tool

Which securityContext settings would you require for every application container, and why? Mid

Each closes a real hole if someone gets code running in the container:

And never privileged: true for an app. Verify from /proc/1/status (Uid, CapEff, NoNewPrivs, Seccomp). The failures tell you what to fix: Permission denied (wrong uid) vs Read-only file system (needs a volume). That block is exactly the restricted Pod Security Standard.

Also asked: A container fails after you enforced readOnlyRootFilesystem and runAsNonRoot. How do you help the team? · What is a Linux capability, and why drop them all? · What does allowPrivilegeEscalation: false do?

Learn it: 17.37 securityContext: the fields worth setting every time

What are the Pod Security Standards, and how would you roll out restricted without breaking teams? Mid

Pod Security Admission is built into the API server and checks pods against three levels, set per namespace with labels:

Three modes: enforce rejects, warn tells the client, audit logs. Roll out in steps:

  1. enforce: baseline + warn: restricted (and audit: restricted) - nothing breaks, everyone sees what would.
  2. Preview: kubectl label ns shop pod-security.kubernetes.io/enforce=restricted --dry-run=server lists the violating pods.
  3. Fix the manifests, then enforce restricted; pin enforce-version so upgrades do not change the rules.

The trap: enforce checks pods, so a Deployment is accepted and its ReplicaSet then cannot create pods - which is why warn (checked on workloads) matters. Running pods are never evicted, only their replacements rejected.

Also asked: What are the limits of Pod Security Admission, and what would you add on top? · Why does a Deployment get accepted in a restricted namespace and still create no pods? · What is the difference between the enforce, warn and audit modes?

Learn it: 17.39 Pod Security Standards and Pod Security Admission

How would you manage application secrets on a Kubernetes platform in a regulated environment? Mid

First, who can read a Secret today - four layers:

  1. RBAC - get, list and watch on secrets (list returns full objects), and anyone who can create pods in the namespace can mount any secret there.
  2. etcd - without encryption at rest, secrets are stored as plain base64, in every backup.
  3. The node - root reads mounted secrets from tmpfs.
  4. The process - env vars leak through /proc/PID/environ, crash dumps and debug endpoints.

Then the controls:

Also asked: Are Kubernetes Secrets secure? · How do you enable encryption at rest for Secrets, and verify it works? · Why is permission to create pods also access to secrets?

Learn it: 17.41 Secrets, and why base64 is not a security measure

How do you give a new person access to a kubeadm cluster with a client certificate? Mid

Kubernetes has no User objects: the certificate is the user. The API server trusts certificates signed by the cluster CA; CN becomes the username, each O a group.

  1. Key and request: openssl genrsa -out jane.key 2048, openssl req -new -key jane.key -out jane.csr -subj "/CN=jane/O=dev".
  2. Submit a CertificateSigningRequest object: request = base64 -w0 jane.csr, signerName: kubernetes.io/kube-apiserver-client, usages: [client auth], a short expirationSeconds.
  3. Check the subject (kubectl describe csr jane), then kubectl certificate approve jane; the controller-manager issues it.
  4. Extract: kubectl get csr jane -o jsonpath='{.status.certificate}' | base64 -d > jane.crt.
  5. kubeconfig: kubectl config set-credentials jane --client-key=jane.key --client-certificate=jane.crt --embed-certs=true, plus set-context.
  6. Authorize: a RoleBinding, preferably to the group (--group=dev).
  7. Test with kubectl --context jane auth whoami and get pods: 401 = certificate problem, 403 = RBAC.

Certificates cannot be revoked, so keep them short-lived - and never O=system:masters.

Also asked: A user's kubectl returns errors. How do you tell whether it is the certificate or RBAC? · Why are client certificates risky for human access in production? · What does a CSR stuck in Approved without Issued tell you?

Learn it: 17.43 Users are certificates: the CSR flow end to end

Practise these answers with flashcards and labs Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.