OnCallReady

Lesson 17.3 · Kubernetes: Scheduling, Health & Security · 17 min read

QoS classes, oom_score_adj, and who gets evicted first

In plain words

Imagine a lifeboat that's getting too heavy. The captain has a rule for who gets off first: passengers who never bought a ticket go first; then passengers who brought more luggage than their ticket allowed, starting with whoever is most over; passengers who bought the exact seat they use are the last to go.

That's QoS. BestEffort pods (no requests or limits at all) are the ticketless passengers. Burstable pods using more than they requested are next. Guaranteed pods (limits equal to requests for CPU and memory, in every container) can never be over their request, so they go last. The kubelet evicts pods when the node's memory.available drops below the threshold, and it also sets each container's oom_score_adj so the kernel picks the same order if it has to act first.

Three classes, computed, never declared

The problem. A node runs low on memory and something has to go. Which pod dies first is decided by a label you never wrote: the pod's QoS class. Knowing the rule lets you protect the payment path and sacrifice the debug pod.

What you need to know already: requests and limits (17.1), the kernel OOM killer and oom_score_adj (5.7), cgroup OOM and exit 137 (5.11).

QoS ("quality of service") class = a label Kubernetes puts on every pod that says how safe it is when memory runs out. You do not set a QoS class. The apiserver computes it at pod creation from the resources of every container (init containers included) and writes it into status.qosClass:

classrule
Guaranteedevery container has cpu AND memory limits, and requests equal to them
Burstablenot Guaranteed, but at least one container has a cpu or memory request or limit
BestEffortno container has any request or limit
# an illustration: pods of each class (the QoS mission builds them)
$ kubectl get pods -n shop -o custom-columns=NAME:.metadata.name,QOS:.status.qosClass
NAME                 QOS
api-7c9f4-k2x8p      Guaranteed
web-5d8f-2kq9x       Burstable
debug-shell          BestEffort

-o custom-columns=HEADER:path,... (15.38) prints only the fields you name: here the column NAME from .metadata.name and QOS from .status.qosClass.

The traps in the rule:

# Guaranteed - limits only: requests are defaulted to the limits
resources:
  limits: {cpu: 500m, memory: 256Mi}
# Burstable - cpu and memory set, but not equal
resources:
  requests: {cpu: 500m, memory: 256Mi}
  limits:   {cpu: "1",  memory: 256Mi}
# Burstable - memory Guaranteed-style, but no cpu at all
resources:
  requests: {memory: 256Mi}
  limits:   {memory: 256Mi}

A two-container pod is Guaranteed only if both containers are. One sidecar without resources (a log shipper someone added last month) turns the whole pod Burstable. The class is immutable for the pod's life: change resources and you get new pods with a newly computed class.

describe pod prints it near the bottom:

QoS Class:                   Burstable
Node-Selectors:              <none>
Tolerations:                 node.kubernetes.io/not-ready:NoExecute op=Exists for 300s
                             node.kubernetes.io/unreachable:NoExecute op=Exists for 300s

What the class changes: two different killers

There are two completely different ways a pod dies for lack of memory. Keep them apart - they have different symptoms, different evidence and different fixes.

1. The container exceeds its own memory limit -> cgroup OOM kill. The kernel kills a process inside that container's cgroup. The pod stays, the container restarts: Last State: Terminated, Reason: OOMKilled, Exit Code: 137. QoS does not matter here; only the limit does.

2. The node runs out of memory -> the kubelet evicts pods. To evict = the kubelet stops a pod on purpose and marks it failed, to save the node. The kubelet watches memory.available on the node. Below the eviction threshold (hard default memory.available<100Mi; also nodefs.available<10% = the node's main disk, imagefs.available<15% = the disk holding container images, nodefs.inodesFree<5% = free inodes, 4.13) it picks pods and evicts them: the pod object ends up Failed with reason Evicted, and its controller creates a replacement elsewhere.

# an illustration: pods of each QoS class (the QoS mission builds them)
kubectl get pods -n batch
NAME                     READY   STATUS    RESTARTS   AGE
cruncher-6f4d8-9zt2m     0/1     Evicted   0          14m
cruncher-6f4d8-xw7kq     1/1     Running   0          40s
kubectl describe pod cruncher-6f4d8-9zt2m -n batch | grep -A2 Status
Status:           Failed
Reason:           Evicted
Message:          The node was low on resource: memory. Threshold quantity: 100Mi, available: 81420Ki. Container cruncher was using 3520112Ki, request is 0, has larger consumption of memory.
# an illustration: pods of each QoS class (the QoS mission builds them)
kubectl describe node worker-2 | tail -4
  Warning  EvictionThresholdMet       2m   kubelet  Attempting to reclaim memory
  Normal   NodeHasInsufficientMemory  2m   kubelet  Node worker-2 status is now: NodeHasInsufficientMemory

Two things to read above: kubectl describe pod ... | grep -A2 Status keeps the Status line and the 2 after it; the Message: says which threshold was crossed and that this container used 3.4 GiB with a request of 0.

While the node is under pressure it has condition MemoryPressure=True (a condition = a true/false health flag the kubelet reports on the node) and the taint node.kubernetes.io/memory-pressure:NoSchedule - no new BestEffort pods land there until it recovers.

The eviction order

The kubelet ranks candidates by, in order:

  1. Does the pod's usage exceed its requests? Pods over their requests go first.
  2. Pod priority (a number set through a PriorityClass, lesson 17.18): lower first.
  3. How far over its requests the pod's usage is: most over first.

This is where QoS shows up in practice. A BestEffort pod requests 0, so any usage at all is "over requests" - it is always in the first group. A Guaranteed pod's usage can never exceed its requests (its limit equals its request, and the cgroup stops it there), so it is evicted only when the node is truly out of options (system daemons eating the memory). Burstable pods are in between: evicted if they are running above their requests.

BestEffort is evicted first, Burstable over its requests next, Guaranteed last. That is the interview answer - and the reason behind it is the ranking above, not a hard-coded list.

Evicted pods do not come back by themselves: the controller replaces them. A bare pod (no Deployment) that gets evicted is gone. kubectl get pods keeps showing Evicted pods until something deletes them (the pod GC does so once there are more than 12500 terminated pods by default - in practice, you clean them up: kubectl delete pods -n batch --field-selector=status.phase=Failed - --field-selector picks objects by a field value instead of by label).

oom_score_adj: the same idea, in the kernel

Chapter 5 (5.7) showed the kernel's OOM killer choosing its victim by oom_score, and you protected sshd with echo -1000 > /proc/PID/oom_score_adj. The kubelet does exactly that for every container, by QoS class, so that if the node runs out of memory before the kubelet reacts, the kernel picks the right victim too:

processoom_score_adj
kubelet, container runtime-999
Guaranteed containers-997
Burstable containersmin(max(2, 1000 - 1000 * memoryRequest / nodeMemoryCapacity), 999)
BestEffort containers1000

So a Burstable pod that requests half the node's memory gets ~500, one that requests 64Mi on a 4Gi node gets ~985 - almost as expendable as BestEffort. You can read it from inside:

# an illustration: pods of each QoS class (the QoS mission builds them)
kubectl exec api-7c9f4-k2x8p -- cat /proc/1/oom_score_adj
-997
kubectl exec debug-shell -- cat /proc/1/oom_score_adj
1000
kubectl exec web-5d8f-2kq9x -n shop -- cat /proc/1/oom_score_adj
968

(For web: 128Mi requested of 4005528Ki capacity -> 1000 - 1000*134217728/4101660672 = 967.3 -> 968 after the kubelet's integer arithmetic.)

Choosing a class on purpose

Note what QoS is not: it is not a priority for CPU scheduling between pods (that is cpu.weight, from requests) and not a scheduling priority (that is PriorityClass, lesson 17.18). It is a label for "how the kubelet and the kernel treat you under memory pressure".

What you can now do

Why it helps

When a node runs low on memory, the order in which pods die is decided by this. On a platform team you'll see a batch job with no requests get the whole node into MemoryPressure, and the evicted pods include a team's API because it was Burstable and slightly over its request. Knowing the ranking lets you explain it and fix it: requests that reflect usage, Guaranteed for the critical path.

It also separates two failures that look alike in dashboards: a container OOMKilled for crossing its own limit (QoS irrelevant) and a pod Evicted because the node ran out (QoS decides). And "who gets evicted first?" is a standard interview question where the real answer is the ranking, not a memorised list.

Commands in this lesson

kubectl

FAQ

Can I set the QoS class myself?

No, it's computed by the API server at pod creation from every container's resources and stored in status.qosClass. Guaranteed needs CPU and memory limits in every container with requests equal to them; BestEffort means no container has any request or limit; everything else is Burstable. Changing resources creates new pods with a newly computed class.

Why is my pod Burstable when I set requests equal to limits?

Usually because one container doesn't follow the rule: a sidecar with no resources, or a container that sets memory but not CPU. Guaranteed requires both CPU and memory, with equal requests and limits, in every container, init containers included. k get pod -o jsonpath='{.status.qosClass}' and check each container's resources.

What's the difference between OOMKilled and Evicted?

OOMKilled: the container exceeded its own memory limit, the kernel killed a process in its cgroup, and the container restarts in the same pod (Last State: OOMKilled, exit 137). Evicted: the node ran low on memory (or disk), the kubelet chose the pod by the eviction ranking, and the pod ends up Failed with reason Evicted; its controller creates a replacement elsewhere.

Do Evicted pods get cleaned up automatically?

Only when the pod garbage collector's threshold is reached, which by default is 12500 terminated pods, so in practice they stay listed. Their controllers have already created replacements. Clean them up with k delete pods -n <ns> --field-selector=status.phase=Failed. A bare pod without a controller that gets evicted is simply gone.

Does QoS affect CPU priority between pods?

No. CPU sharing under contention comes from cpu.weight, which is derived from the CPU request. Scheduling priority and preemption come from PriorityClass. QoS is only about how the kubelet and the kernel treat the pod under memory pressure: eviction order and oom_score_adj.

In an interview Junior

What are the Kubernetes QoS classes, and which pods get evicted first?

The QoS class is computed from every container's resources (never declared) and written to status.qosClass:

When a node runs low (below the eviction threshold, e.g. memory.available<100Mi), the kubelet evicts pods, ranked by: is usage above the requests, then priority, then how far above. So BestEffort (request 0, always "above") goes first, Burstable over its requests next, Guaranteed last. The kubelet also sets oom_score_adj by class (Guaranteed -997, BestEffort 1000) so the kernel agrees if it acts first.

Keep that apart from a container hitting its own memory limit: a cgroup OOM kill, the container restarts with exit 137, whatever the class.

Also asked: How do you make a critical service less likely to be evicted? · What is the difference between an eviction and an OOM kill? · Why can one sidecar change a pod's QoS class?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.