OnCallReady

Lesson 17.1 · Kubernetes: Scheduling, Health & Security · 24 min read

Requests and limits: what each one actually controls

In plain words

Think of booking tables at a restaurant. When you book, you say "we're four people" and the restaurant holds four seats for you, even if only two of you turn up. That booking is how the host decides whether another group fits tonight. Separately, the restaurant has a rule: "no table may take more than six chairs", and a waiter physically stops you taking a seventh.

A request is the booking: the scheduler adds up requests on each node and compares them with allocatable, ignoring how busy the node really is. A limit is the waiter: the kernel's cgroup enforces it, throttling CPU above cpu.max and OOM-killing a container that crosses memory.max. kubectl describe node shows the bookings; kubectl top shows who actually turned up.

Two numbers, two different machines

The problem. A pod sits Pending with "Insufficient cpu" while the nodes are almost idle. Another pod is killed with exit 137 although "the node has plenty of memory". Both come from two numbers in the pod spec that most people set by copy-paste. This lesson explains what each number really does.

What you need to know already: pods, Deployments and kubectl (15.1-15.16), the scheduler and the kubelet (15.5, 15.7), cgroups, MemoryMax=/CPUQuota= and exit 137 (2.26, 5.11), container memory/CPU limits in Docker (11.10).

A few words first, in plain terms:

The lessons use chapter 15's speed kit. If this shell has no k (you jumped in here), define it now (alias = a short name for a longer command):

$ alias k=kubectl

Every container can declare two numbers per resource:

resources:
  requests:          # what the SCHEDULER reserves for you
    cpu: 250m
    memory: 256Mi
  limits:            # what the KERNEL (cgroup) will never let you exceed
    cpu: "1"
    memory: 512Mi

The units, in one line each (details at the end of the lesson): 250m = 250 millicores = a quarter of one CPU core; "1" = one whole core; 256Mi = 256 mebibytes (1 Mi = 1024 x 1024 bytes).

They are read by different components at different times, and confusing them is the root of half the "Kubernetes is slow / Kubernetes killed my app" tickets you will ever see:

requestslimits
read bykube-scheduler (placement), kubelet (eviction ranking, cgroup weights)kubelet -> container runtime -> cgroup
whenonce, when the pod is scheduledcontinuously, every CFS period, every page the app touches
cpu meansa share of the CPU when it is contendeda hard quota: exceed it and you are throttled
memory means"reserve this much on the node"a hard ceiling: exceed it and the kernel OOM-kills you
if absent0 - you reserve nothingunbounded - you can use the whole node

(CFS period: the Linux CPU scheduler - CFS, the "Completely Fair Scheduler" - hands out CPU time in slices of 100 ms; a CPU limit is "at most X ms of CPU per 100 ms slice". Throttled = paused until the next slice because the quota ran out.)

The scheduler never looks at limits, and never looks at actual usage. It only adds up the requests of the pods on a node and compares them with the node's allocatable.

The scheduler's ledger: allocatable minus requests

A node reports two sets of numbers. kubectl describe node worker-1 prints everything the cluster knows about that node; we pipe it into grep (7.1) to keep only two blocks: -E = extended regex, '^(Capacity|Allocatable)' = lines starting with either word, -A7 = plus the 7 lines After each match.

$ kubectl describe node worker-1 | grep -A7 -E '^(Capacity|Allocatable)'
Capacity:
  cpu:                2
  ephemeral-storage:  19495744Ki
  memory:             4005528Ki
  pods:               110
Allocatable:
  cpu:                2
  ephemeral-storage:  17967521792
  memory:             3903128Ki
  pods:               110

Capacity is the hardware. Allocatable is capacity minus what the kubelet keeps back for the OS and itself (--kube-reserved, --system-reserved: kubelet settings that set memory/CPU aside for system daemons) and the eviction threshold (the "too little free memory left" line where the kubelet starts removing pods - memory.available<100Mi by default, which is exactly the ~100Mi difference here). Pods can only be placed into allocatable. The other rows: ephemeral-storage = local disk for container files and logs, pods: 110 = the most pods this node accepts.

Further down describe node is the ledger itself:

Non-terminated Pods:          (6 in total)
  Namespace    Name                 CPU Requests  CPU Limits  Memory Requests  Memory Limits  Age
  ---------    ----                 ------------  ----------  ---------------  -------------  ---
  kube-system  calico-node-tzf78    250m (12%)    0 (0%)      0 (0%)           0 (0%)         12d
  kube-system  coredns-d6qs-lmrzs   100m (5%)     0 (0%)      70Mi (1%)        170Mi (4%)     12d
  kube-system  kube-proxy-r47xw     0 (0%)        0 (0%)      0 (0%)           0 (0%)         12d
  shop         web-5d8f-2kq9x       100m (5%)     500m (25%)  128Mi (3%)       256Mi (6%)     2h
Allocated resources:
  (Total limits may be over 100 percent, i.e., overcommitted.)
  Resource            Requests     Limits
  --------            --------     ------
  cpu                 450m (22%)   500m (25%)
  memory              198Mi (5%)   426Mi (11%)

Read it as a bank statement. Requests of 450m out of 2 cores are "spent" - a new pod asking for 1600m does not fit, even if the node's real CPU usage is 3%. That is why you see this, on a cluster whose CPU graphs are flat:

Warning  FailedScheduling  default-scheduler  0/3 nodes are available: 1 node(s) had untolerated taint {node-role.kubernetes.io/control-plane: }, 2 Insufficient cpu. preemption: 0/3 nodes are available: 1 Preemption is not helpful for scheduling, 2 No preemption victims found for incoming pod.

Read that event piece by piece: FailedScheduling = the scheduler could not place the pod; 0/3 nodes are available = none of the 3 nodes fits; then one reason per node group - one node has a taint (a "keep out" mark, lesson 17.13) it does not tolerate, two have Insufficient cpu. The preemption: part says the scheduler also could not make room by removing lower-priority pods (lesson 17.18).

And the opposite: a pod with no requests reserves nothing, so the scheduler will happily stack twenty of them on a node that then runs out of real memory. "Insufficient cpu" means requested, never used.

(Total limits may be over 100 percent, i.e., overcommitted.) is kubectl telling you the other half: limits are not reserved at all. The sum of limits on a node routinely exceeds 100% - that is overcommit, and it is normal. It only hurts when enough pods actually use their limits at the same time.

Compare with what is really used right now. kubectl top node shows live usage; it needs metrics-server, a small add-on that asks every kubelet for CPU and memory usage (sampled every 15s) and serves it to kubectl:

$ kubectl top node
NAME       CPU(cores)   CPU(%)   MEMORY(bytes)   MEMORY(%)
cp-1       173m         9%       1388Mi          36%
worker-1   67m          3%       656Mi           17%
worker-2   114m         6%       1019Mi          27%

top percentages are against allocatable, not against requests. The gap between "Requests 22%" and "CPU 3%" is money: requests you pay for (nodes) and do not use. Right-sizing requests is one of the first things a platform team does.

What the kubelet does with the numbers: cgroups

The kubelet turns each container's resources into cgroup v2 settings (the same files you read under /sys/fs/cgroup in 5.11 and 11.10). You can read them from inside the container: kubectl exec <pod> -n <ns> -- <command> runs a command in the container (15.42), and there - thanks to a private cgroup namespace (containerd's default: the container sees only its own branch of the cgroup tree) - /sys/fs/cgroup is the container's own cgroup:

# shop/web with the resources above (an illustration; the next mission sets them for real)
kubectl exec web-5d8f-2kq9x -n shop -- cat /sys/fs/cgroup/cpu.max
50000 100000
kubectl exec web-5d8f-2kq9x -n shop -- cat /sys/fs/cgroup/cpu.weight
4
kubectl exec web-5d8f-2kq9x -n shop -- cat /sys/fs/cgroup/memory.max
268435456

This is the same cgroup machinery as Ch 2's (lesson 2.26) MemoryMax= and CPUQuota= on a systemd unit - systemd-run -p MemoryMax=200M killed memhog with 137 the same way a container limit does. Kubernetes just writes the files for you.

Units, and the mistakes they cause

cpu:     1 = 1000m = one core (vCPU).  0.5 = 500m.  "100m" = a tenth of a core.
memory:  Mi/Gi = powers of 1024.  M/G = powers of 1000.  128Mi = 134217728 bytes.
memory: 128m      # 0.128 BYTES - a lowercase m is milli, not mega
memory: 1G        # 1000000000 bytes, ~7% less than 1Gi
cpu: 1000         # one thousand cores

The apiserver accepts all three - they are valid quantities - and the pod either never schedules (1000 cores) or is killed immediately (0.128 bytes). Always use Mi/Gi for memory and m for CPU.

Two validation errors you will hit when editing by hand:

The Pod "web" is invalid: spec.containers[0].resources.requests: Invalid value: "2": must be less than or equal to cpu limit of 1
The Deployment "web" is invalid: spec.template.spec.containers[0].resources.limits[memory]: Invalid value: "512mb": must be a valid quantity

A request may never exceed its limit. And if you set only a limit, the apiserver copies it into the request - which quietly makes that pod Guaranteed (next lesson) and reserves the full limit on the node.

Setting them

# shop/web with the resources above (an illustration; the next mission sets them for real)
kubectl set resources deploy/web -n shop -c web --requests=cpu=100m,memory=128Mi --limits=cpu=500m,memory=256Mi
deployment.apps/web resource requirements updated

kubectl set resources edits the resources in a workload's pod template: deploy/web = the Deployment called web; -n shop = in namespace shop; -c web = only the container named web; --requests= / --limits= = comma-separated resource=amount pairs. The alternative is editing the YAML (kubectl edit, 15.40) - same result.

Changing resources changes the pod template, so it is a rollout: new pods, new cgroups. (In-place pod resize - beta in 1.33, GA in 1.35 - resizes a running container's cgroup without a restart via the pod's resize subresource, but changing a Deployment's template still rolls new pods.)

Pod-level totals: the scheduler adds up the containers. Init containers run one at a time, so a pod's effective request is max(largest init container, sum of app containers) - plus sidecars (init containers with restartPolicy: Always), which run alongside and are added to the sum.

What you can now do

Why it helps

Half of "Kubernetes is slow / killed my app" tickets come from mixing these up. Pods Pending with "Insufficient cpu" on a cluster whose CPU graphs are flat: requests, not usage. Twenty pods with no requests packed onto one node that then runs out of memory: nothing reserved. memory: 128m in a YAML file, which is 0.128 bytes: a unit typo that passes validation.

As a platform engineer you'll also own the money side: the gap between requested and used CPU is nodes you pay for and don't use, and right-sizing requests is one of the first cost projects on any team. Reading the cgroup files (cpu.max, cpu.weight, memory.max) links this straight back to the systemd cgroups you already know from Ch 2 and Ch 5.

Commands in this lesson

alias kubectl

FAQ

Does the scheduler look at how much CPU a node is really using?

No. It only compares the sum of requests of the pods on a node with the node's allocatable. Actual usage and limits are ignored. So a node at 3% real CPU can reject a pod ("Insufficient cpu"), and a node that's nearly out of memory can accept one, if the existing pods under-request.

What happens if I set only a limit?

The API server copies the limit into the request. The pod then reserves the full limit on the node, and if every container has CPU and memory limits set this way, the pod becomes Guaranteed QoS. That's fine when intended, and a quiet waste of capacity when not.

What's the difference between Mi and M, and why is 128m wrong?

Mi and Gi are powers of 1024; M and G are powers of 1000, so 1G is about 7% less than 1Gi. A lowercase m means milli, so memory: 128m is 0.128 bytes. The API server accepts it as a valid quantity and the container is killed immediately. Use Mi/Gi for memory and m for CPU.

Is it bad that the total limits on a node are over 100%?

Not by itself. Limits aren't reserved, so overcommitting them is normal: most pods don't hit their limits at the same time. It hurts when enough of them do. Memory overcommit is riskier than CPU, because CPU contention only slows pods down while memory contention leads to evictions and OOM kills.

Is cpu.weight the same as a CPU limit?

No. cpu.weight comes from the CPU request and only matters when the node's CPU is contended: each cgroup then gets CPU in proportion to its weight. Uncontended, a pod can burst above its request. cpu.max comes from the limit and is a hard quota per 100 ms period, enforced even if the node is idle.

In an interview Junior

Explain resource requests and limits in Kubernetes.

Two numbers per resource, read by different components:

Consequences: Insufficient cpu on an idle cluster means requested, not used; a pod without requests is invisible to the scheduler and gets overpacked and evicted first; limits are not reserved, so they may add up past 100% (overcommit). Units: 250m = a quarter core, 256Mi = mebibytes - 128m memory is 0.128 bytes. kubectl describe node shows the ledger.

Also asked: The cluster is only 20% utilised, yet new pods do not schedule. Why? · How do you choose request and limit values for a service? · What happens if you set only a limit and no request?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.