OnCallReady

Lesson 19.2 · CKA Exam Drilling · 11 min read

The five domains and the tasks they turn into

In plain words

Imagine a school exam in five subjects, where the teacher tells you in advance how many marks each subject is worth: fixing broken things 30%, building and running the school 25%, the corridors and phones 20%, organising the classes 15%, and the storage rooms 10%. And every question in each subject looks like one of a handful of patterns you've seen before.

The CKA's five domains are Troubleshooting (30%), Cluster Architecture, Installation & Configuration (25%), Services & Networking (20%), Workloads & Scheduling (15%) and Storage (10%). Each turns into a small set of task shapes: "Deployment has no ready pods", "RBAC for a ServiceAccount", "NetworkPolicy allow-from", "PV plus PVC plus pod". Recognise the shape and you know the first command before you finish reading.

Five domains, weighted

The problem. Two hours is not enough to be good at everything. The exam publishes how much each area is worth; spend your practice where the points are.

What you need to know already: the exam format (19.1); every command in the tables below was taught in chapters 15-18 - the lesson numbers are given where it helps.

The CNCF curriculum (the official list of what the exam covers; v1.35, github.com/cncf/curriculum, checked September 2026):

domainweightcompetencies
Troubleshooting30%clusters and nodes; cluster components; resource usage; container output streams; services and networking
Cluster Architecture, Installation & Configuration25%RBAC; prepare infrastructure; kubeadm clusters; cluster lifecycle (upgrades, etcd); Helm and Kustomize (packaging and templating tools for YAML); extension interfaces (CNI, CSI - the storage plugin interface, CRI); CRDs and operators
Services & Networking20%pod connectivity; NetworkPolicies; ClusterIP/NodePort/LoadBalancer and endpoints; Gateway API; Ingress; CoreDNS
Workloads & Scheduling15%Deployments, rolling updates, rollbacks; ConfigMaps and Secrets; workload autoscaling; self-healing primitives; admission and scheduling (limits, affinity...)
Storage10%StorageClasses and dynamic provisioning; volume types, access modes, reclaim policies; PVs and PVCs

Troubleshooting plus architecture is more than half the exam. That is where the drill rounds in this chapter are concentrated.

The curriculum changed in February 2025 (Helm/Kustomize, Gateway API, CRDs/operators, extension interfaces were added; the weights moved to the ones above). Check the repository once before you book - the weights are stable, the competency wording moves a little each version.

The task shapes

Across practice exams the tasks come in a small number of shapes. Recognise the shape and you know the first command before you finish reading.

Troubleshooting (30%)

shapefirst commands
"Deployment X has no ready pods, fix it"k get pods -n NS (STATUS), k describe pod (Events, Last State), k logs --previous
"Service exists, pods run, no traffic"k get endpointslices -n NS, --show-labels vs selector, targetPort vs containerPort
"Node X is NotReady"k describe node (Conditions), ssh, systemctl status kubelet containerd, journalctl -u kubelet
"New pods do not start / kubectl refuses"k get pods -n kube-system, on the control plane crictl ps -a, crictl logs, the manifest in /etc/kubernetes/manifests
"Write the ERROR lines of container Y to a file"k logs POD -c Y | grep ... > file
"Which pod uses most CPU/memory"k top pod -A --sort-by=memory
# an illustration: exam-style task state (the mock tasks build it)
k get pods -n shop
NAME                   READY   STATUS                       RESTARTS      AGE
api-5d8f9c7b6-2kq9x    0/1     CreateContainerConfigError   0             2m
web-7c9d8f6b5-8xz2m    0/1     CrashLoopBackOff             4 (31s ago)   2m
k describe pod -n shop api-5d8f9c7b6-2kq9x | tail -3
  Warning  Failed     12s (x9 over 2m)  kubelet   Error: couldn't find key log-level in ConfigMap shop/api-config

The STATUS column already told you which chapter of the failure catalogue you are in; the last Event names the object to fix.

Cluster architecture (25%)

shapefastest route
RBAC for a ServiceAccount / userk create sa, k create role --verb --resource, k create rolebinding --role --serviceaccount=NS:NAME, prove with k auth can-i ... --as
Cluster-scoped read for a userk create clusterrole + clusterrolebinding --user
etcd backupetcdctl snapshot save with the 3 TLS files from etcd.yaml, etcdutl snapshot status
etcd restoreetcdutl snapshot restore --data-dir NEW, change the etcd-data hostPath
Upgrade control plane / workerrepository per minor, kubeadm first, kubeadm upgrade plan/apply or upgrade node, drain, kubelet+kubectl, restart, uncordon
Maintenancek drain --ignore-daemonsets --delete-emptydir-data, k uncordon
Contexts, certificates, joink config get-contexts -o name, openssl x509 -noout -enddate, kubeadm token create --print-join-command
# an illustration: exam-style task state (the mock tasks build it)
k auth can-i list secrets -n team-a --as=system:serviceaccount:team-a:reader
no
k auth can-i list pods -n team-a --as=system:serviceaccount:team-a:reader
yes

Always run both: the one that must say yes, and one that must say no.

Helm, Kustomize and CRDs/operators are in the curriculum. Helm is a package manager for Kubernetes: a chart is a package of YAML templates, --set key=value fills in its settings (values), and an installed chart is a release. Kustomize (built into kubectl as k apply -k DIR) layers small patches over plain YAML. Expect "install this chart with these values" (helm install NAME REPO/CHART -n NS --create-namespace --set k=v) or "apply this kustomization" (k apply -k DIR), and "list the CRDs of operator X / create a custom resource" (CRDs and operators: 17.30, 17.41). (simulator) The lab has no helm binary and no kubectl apply -k yet; practise them on KillerCoda (a free browser-based practice environment).

Later (Ch 25): Helm gets its own lessons, with charts you write yourself.

Services & networking (20%)

shapefastest route
Expose on a NodePortk expose deploy X --type=NodePort --port --target-port, then patch nodePort
NetworkPolicy allow-fromcopy the docs example; podSelector = the protected pods; peers under from; the POD port
Ingressk create ingress NAME --class=nginx --rule="host/path*=svc:port"
Gateway APIHTTPRoute from the gateway-api docs: parentRefs, hostnames, matches, backendRefs
DNSNAME.NS.svc.cluster.local, pod A record a-b-c-d.NS.pod.cluster.local, nslookup from a pod

Workloads & scheduling (15%)

shapefastest route
Rollout / rollbackk set image, k rollout history --revision=N, k rollout undo --to-revision=N
ConfigMap/Secret into a podk create cm/secret --from-literal, k run $do, add volume / env valueFrom
Resourcesk set resources deploy/X --requests=... --limits=...
Taints, affinity, nodeSelector`k describe nodegrep Taints, toleration + selector; nodeAffinity from k explain --recursive`
Probes, sidecaredit the container: readinessProbe/livenessProbe; initContainers with restartPolicy: Always
CronJob / Jobk create cronjob --schedule $do, add history limits and activeDeadlineSeconds; k create job --from=cronjob/X
HPAk autoscale deploy X --min --max --cpu=70%

Storage (10%)

shapefastest route
PV + PVC + poddocs "Configure a Pod to Use a PersistentVolume for Storage": all three objects on one page
StorageClass, default classstorageclass.kubernetes.io/is-default-class annotation - exactly one default
Expand a claimallowVolumeExpansion on the class, raise spec.resources.requests.storage
# an illustration: exam-style task state (the mock tasks build it)
k get pvc -n data
NAME      STATUS    VOLUME   CAPACITY   ACCESS MODES   STORAGECLASS   AGE
db-data   Pending                                      manual         40s
k get pv
NAME    CAPACITY   ACCESS MODES   RECLAIM POLICY   STATUS      CLAIM   STORAGECLASS   AGE
db-pv   1Gi        RWX            Retain           Available           manual         2m

Pending with an Available PV right there: the access modes differ (claim RWO, volume RWX). A claim binds only if class, access modes and size all fit.

What this means for practice

Each shape above is one drill in this chapter. The drills do not teach you the shape - chapters 15-18 did - they make it fast. Work through them by weight: troubleshooting and RBAC first, then networking, then the rest; and redo any drill whose round took longer than its budget.

Why it helps

The weights tell you where to spend practice time: troubleshooting plus architecture is more than half the exam, so those drills come first, and storage, while important, is a tenth. Studying by weight instead of by chapter order is the most efficient preparation you can do in the last weeks.

The task shapes are what make you fast: when "Service exists, pods run, no traffic" appears, your hands already type k get endpointslices and compare labels. It's the same pattern matching that makes an experienced on-call engineer quick. The February 2025 curriculum change (Helm and Kustomize - packaging and patching tools for YAML - Gateway API, CRDs and operators) also means older courses and blog posts miss topics; knowing the current list keeps you from being surprised.

FAQ

Where do I find the official curriculum?

In the CNCF curriculum repository on GitHub (github.com/cncf/curriculum), per Kubernetes version. It lists the five domains, their weights and the competencies in each. The weights have been stable since the February 2025 update; the wording of competencies moves a little with each version, so check it once before booking.

What changed in the 2025 curriculum?

Helm and Kustomize, Gateway API, CRDs and operators, and extension interfaces (CNI, CSI, CRI) were added, and the weights shifted to the current ones, with troubleshooting at 30%. Older courses, practice sets and blog posts written before 2025 don't cover these, so expect tasks like installing a chart with values or creating an HTTPRoute.

Why is my PVC Pending when there's an Available PV?

A claim binds only if the StorageClass name, the access modes and the size all fit. The classic exam version: the claim asks for ReadWriteOnce and the PV offers only ReadWriteMany, or the class names differ. Compare k get pvc and k get pv side by side; the ACCESS MODES and STORAGECLASS columns usually show the mismatch.

Do I need to verify RBAC with a "no" as well as a "yes"?

Yes. A yes for the permission the task asks for proves you granted enough; a no for something it shouldn't have (like list secrets) proves you didn't grant too much. Tasks often say "only" or "exactly", and over-granting fails those checks. k auth can-i ... --as=system:serviceaccount:NS:NAME does both in seconds.

Which domain should I drill first?

Troubleshooting, then cluster architecture (RBAC, etcd, upgrades, kubeadm), then networking, then workloads and storage. That's the order of points. Within each, redo any drill whose round took longer than its budget, since speed, not knowledge, is usually what's missing by this stage.

In an interview Junior

A Deployment has no ready pods. What are your first commands?

  1. k get pods -n NS - the STATUS column says which part of the failure catalogue you are in: Pending, ImagePullBackOff, CrashLoopBackOff, CreateContainerConfigError, Running but 0/1.
  2. k describe pod POD -n NS - the Events at the bottom and Last State (exit code, reason). The last event names the object to fix: the scheduler's reason, the registry's answer, a missing ConfigMap key, a failing probe.
  3. k logs POD -n NS --previous - for a crash loop, the run that died.
  4. If no pods exist at all: k describe rs -n NS - quota, Pod Security or a missing ServiceAccount rejects them there.

Then fix in place - k set image, k set env, k set resources, k edit, k rollout undo - rather than deleting and recreating, and verify with k get pods until 1/1 Running with no new restarts.

Also asked: A Service exists and its pods are running, but there is no traffic. What do you check? · How would you give a ServiceAccount read access to pods in one namespace, and prove it? · Which domains carry the most weight in the CKA, and how does that shape practice?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.