Three classes, computed, never declared
The problem. A node runs low on memory and something has to go. Which pod dies first is decided by a label you never wrote: the pod's QoS class. Knowing the rule lets you protect the payment path and sacrifice the debug pod.
What you need to know already: requests and limits (17.1), the kernel OOM killer and oom_score_adj (5.7), cgroup OOM and exit 137 (5.11).
QoS ("quality of service") class = a label Kubernetes puts on every pod that says how safe it is when memory runs out. You do not set a QoS class. The apiserver computes it at pod creation from the resources of every container (init containers included) and writes it into status.qosClass:
| class | rule |
|---|---|
| Guaranteed | every container has cpu AND memory limits, and requests equal to them |
| Burstable | not Guaranteed, but at least one container has a cpu or memory request or limit |
| BestEffort | no container has any request or limit |
# an illustration: pods of each class (the QoS mission builds them)
$ kubectl get pods -n shop -o custom-columns=NAME:.metadata.name,QOS:.status.qosClass
NAME QOS
api-7c9f4-k2x8p Guaranteed
web-5d8f-2kq9x Burstable
debug-shell BestEffort
-o custom-columns=HEADER:path,... (15.38) prints only the fields you name: here the column NAME from .metadata.name and QOS from .status.qosClass.
The traps in the rule:
# Guaranteed - limits only: requests are defaulted to the limits
resources:
limits: {cpu: 500m, memory: 256Mi}
# Burstable - cpu and memory set, but not equal
resources:
requests: {cpu: 500m, memory: 256Mi}
limits: {cpu: "1", memory: 256Mi}
# Burstable - memory Guaranteed-style, but no cpu at all
resources:
requests: {memory: 256Mi}
limits: {memory: 256Mi}
A two-container pod is Guaranteed only if both containers are. One sidecar without resources (a log shipper someone added last month) turns the whole pod Burstable. The class is immutable for the pod's life: change resources and you get new pods with a newly computed class.
describe pod prints it near the bottom:
QoS Class: Burstable
Node-Selectors: <none>
Tolerations: node.kubernetes.io/not-ready:NoExecute op=Exists for 300s
node.kubernetes.io/unreachable:NoExecute op=Exists for 300s
What the class changes: two different killers
There are two completely different ways a pod dies for lack of memory. Keep them apart - they have different symptoms, different evidence and different fixes.
1. The container exceeds its own memory limit -> cgroup OOM kill. The kernel kills a process inside that container's cgroup. The pod stays, the container restarts: Last State: Terminated, Reason: OOMKilled, Exit Code: 137. QoS does not matter here; only the limit does.
2. The node runs out of memory -> the kubelet evicts pods. To evict = the kubelet stops a pod on purpose and marks it failed, to save the node. The kubelet watches memory.available on the node. Below the eviction threshold (hard default memory.available<100Mi; also nodefs.available<10% = the node's main disk, imagefs.available<15% = the disk holding container images, nodefs.inodesFree<5% = free inodes, 4.13) it picks pods and evicts them: the pod object ends up Failed with reason Evicted, and its controller creates a replacement elsewhere.
# an illustration: pods of each QoS class (the QoS mission builds them)
kubectl get pods -n batch
NAME READY STATUS RESTARTS AGE
cruncher-6f4d8-9zt2m 0/1 Evicted 0 14m
cruncher-6f4d8-xw7kq 1/1 Running 0 40s
kubectl describe pod cruncher-6f4d8-9zt2m -n batch | grep -A2 Status
Status: Failed
Reason: Evicted
Message: The node was low on resource: memory. Threshold quantity: 100Mi, available: 81420Ki. Container cruncher was using 3520112Ki, request is 0, has larger consumption of memory.
# an illustration: pods of each QoS class (the QoS mission builds them)
kubectl describe node worker-2 | tail -4
Warning EvictionThresholdMet 2m kubelet Attempting to reclaim memory
Normal NodeHasInsufficientMemory 2m kubelet Node worker-2 status is now: NodeHasInsufficientMemory
Two things to read above: kubectl describe pod ... | grep -A2 Status keeps the Status line and the 2 after it; the Message: says which threshold was crossed and that this container used 3.4 GiB with a request of 0.
While the node is under pressure it has condition MemoryPressure=True (a condition = a true/false health flag the kubelet reports on the node) and the taint node.kubernetes.io/memory-pressure:NoSchedule - no new BestEffort pods land there until it recovers.
The eviction order
The kubelet ranks candidates by, in order:
- Does the pod's usage exceed its requests? Pods over their requests go first.
- Pod priority (a number set through a PriorityClass, lesson 17.18): lower first.
- How far over its requests the pod's usage is: most over first.
This is where QoS shows up in practice. A BestEffort pod requests 0, so any usage at all is "over requests" - it is always in the first group. A Guaranteed pod's usage can never exceed its requests (its limit equals its request, and the cgroup stops it there), so it is evicted only when the node is truly out of options (system daemons eating the memory). Burstable pods are in between: evicted if they are running above their requests.
BestEffort is evicted first, Burstable over its requests next, Guaranteed last. That is the interview answer - and the reason behind it is the ranking above, not a hard-coded list.
Evicted pods do not come back by themselves: the controller replaces them. A bare pod (no Deployment) that gets evicted is gone. kubectl get pods keeps showing Evicted pods until something deletes them (the pod GC does so once there are more than 12500 terminated pods by default - in practice, you clean them up: kubectl delete pods -n batch --field-selector=status.phase=Failed - --field-selector picks objects by a field value instead of by label).
oom_score_adj: the same idea, in the kernel
Chapter 5 (5.7) showed the kernel's OOM killer choosing its victim by oom_score, and you protected sshd with echo -1000 > /proc/PID/oom_score_adj. The kubelet does exactly that for every container, by QoS class, so that if the node runs out of memory before the kubelet reacts, the kernel picks the right victim too:
| process | oom_score_adj |
|---|---|
| kubelet, container runtime | -999 |
| Guaranteed containers | -997 |
| Burstable containers | min(max(2, 1000 - 1000 * memoryRequest / nodeMemoryCapacity), 999) |
| BestEffort containers | 1000 |
So a Burstable pod that requests half the node's memory gets ~500, one that requests 64Mi on a 4Gi node gets ~985 - almost as expendable as BestEffort. You can read it from inside:
# an illustration: pods of each QoS class (the QoS mission builds them)
kubectl exec api-7c9f4-k2x8p -- cat /proc/1/oom_score_adj
-997
kubectl exec debug-shell -- cat /proc/1/oom_score_adj
1000
kubectl exec web-5d8f-2kq9x -n shop -- cat /proc/1/oom_score_adj
968
(For web: 128Mi requested of 4005528Ki capacity -> 1000 - 1000*134217728/4101660672 = 967.3 -> 968 after the kubelet's integer arithmetic.)
Choosing a class on purpose
- Guaranteed for things that must not be evicted and must have predictable latency: databases, the payment path, anything stateful. The price: you pay for the full limit on the node all the time.
- Burstable is the default for most services: request what you normally use, limit memory at what you can tolerate.
- BestEffort only for genuinely disposable work (a debug pod, a re-runnable batch job). Never for anything with users behind it.
Note what QoS is not: it is not a priority for CPU scheduling between pods (that is cpu.weight, from requests) and not a scheduling priority (that is PriorityClass, lesson 17.18). It is a label for "how the kubelet and the kernel treat you under memory pressure".
What you can now do
- Predict a pod's QoS class from its resources, and read it with
custom-columns. - Tell a cgroup OOM kill (container restarts, 137) from a node eviction (pod
Evicted). - Explain the eviction order and the
oom_score_adjvalues behind it.