Two numbers, two different machines
The problem. A pod sits Pending with "Insufficient cpu" while the nodes are almost idle. Another pod is killed with exit 137 although "the node has plenty of memory". Both come from two numbers in the pod spec that most people set by copy-paste. This lesson explains what each number really does.
What you need to know already: pods, Deployments and kubectl (15.1-15.16), the scheduler and the kubelet (15.5, 15.7), cgroups, MemoryMax=/CPUQuota= and exit 137 (2.26, 5.11), container memory/CPU limits in Docker (11.10).
A few words first, in plain terms:
- resource - here: CPU time or memory (RAM) a container uses.
- request - the amount you ask to have reserved for the container.
- limit - the most the container is allowed to use; the kernel enforces it.
- scheduler (
kube-scheduler, 15.5) - the control-plane program that picks a node for each new pod. - kubelet (15.7) - the agent on each node that starts the containers and writes their cgroup settings.
The lessons use chapter 15's speed kit. If this shell has no k (you jumped in here), define it now (alias = a short name for a longer command):
$ alias k=kubectl
Every container can declare two numbers per resource:
resources:
requests: # what the SCHEDULER reserves for you
cpu: 250m
memory: 256Mi
limits: # what the KERNEL (cgroup) will never let you exceed
cpu: "1"
memory: 512Mi
The units, in one line each (details at the end of the lesson): 250m = 250 millicores = a quarter of one CPU core; "1" = one whole core; 256Mi = 256 mebibytes (1 Mi = 1024 x 1024 bytes).
They are read by different components at different times, and confusing them is the root of half the "Kubernetes is slow / Kubernetes killed my app" tickets you will ever see:
| requests | limits | |
|---|---|---|
| read by | kube-scheduler (placement), kubelet (eviction ranking, cgroup weights) | kubelet -> container runtime -> cgroup |
| when | once, when the pod is scheduled | continuously, every CFS period, every page the app touches |
| cpu means | a share of the CPU when it is contended | a hard quota: exceed it and you are throttled |
| memory means | "reserve this much on the node" | a hard ceiling: exceed it and the kernel OOM-kills you |
| if absent | 0 - you reserve nothing | unbounded - you can use the whole node |
(CFS period: the Linux CPU scheduler - CFS, the "Completely Fair Scheduler" - hands out CPU time in slices of 100 ms; a CPU limit is "at most X ms of CPU per 100 ms slice". Throttled = paused until the next slice because the quota ran out.)
The scheduler never looks at limits, and never looks at actual usage. It only adds up the requests of the pods on a node and compares them with the node's allocatable.
The scheduler's ledger: allocatable minus requests
A node reports two sets of numbers. kubectl describe node worker-1 prints everything the cluster knows about that node; we pipe it into grep (7.1) to keep only two blocks: -E = extended regex, '^(Capacity|Allocatable)' = lines starting with either word, -A7 = plus the 7 lines After each match.
$ kubectl describe node worker-1 | grep -A7 -E '^(Capacity|Allocatable)'
Capacity:
cpu: 2
ephemeral-storage: 19495744Ki
memory: 4005528Ki
pods: 110
Allocatable:
cpu: 2
ephemeral-storage: 17967521792
memory: 3903128Ki
pods: 110
Capacity is the hardware. Allocatable is capacity minus what the kubelet keeps back for the OS and itself (--kube-reserved, --system-reserved: kubelet settings that set memory/CPU aside for system daemons) and the eviction threshold (the "too little free memory left" line where the kubelet starts removing pods - memory.available<100Mi by default, which is exactly the ~100Mi difference here). Pods can only be placed into allocatable. The other rows: ephemeral-storage = local disk for container files and logs, pods: 110 = the most pods this node accepts.
Further down describe node is the ledger itself:
Non-terminated Pods: (6 in total)
Namespace Name CPU Requests CPU Limits Memory Requests Memory Limits Age
--------- ---- ------------ ---------- --------------- ------------- ---
kube-system calico-node-tzf78 250m (12%) 0 (0%) 0 (0%) 0 (0%) 12d
kube-system coredns-d6qs-lmrzs 100m (5%) 0 (0%) 70Mi (1%) 170Mi (4%) 12d
kube-system kube-proxy-r47xw 0 (0%) 0 (0%) 0 (0%) 0 (0%) 12d
shop web-5d8f-2kq9x 100m (5%) 500m (25%) 128Mi (3%) 256Mi (6%) 2h
Allocated resources:
(Total limits may be over 100 percent, i.e., overcommitted.)
Resource Requests Limits
-------- -------- ------
cpu 450m (22%) 500m (25%)
memory 198Mi (5%) 426Mi (11%)
Read it as a bank statement. Requests of 450m out of 2 cores are "spent" - a new pod asking for 1600m does not fit, even if the node's real CPU usage is 3%. That is why you see this, on a cluster whose CPU graphs are flat:
Warning FailedScheduling default-scheduler 0/3 nodes are available: 1 node(s) had untolerated taint {node-role.kubernetes.io/control-plane: }, 2 Insufficient cpu. preemption: 0/3 nodes are available: 1 Preemption is not helpful for scheduling, 2 No preemption victims found for incoming pod.
Read that event piece by piece: FailedScheduling = the scheduler could not place the pod; 0/3 nodes are available = none of the 3 nodes fits; then one reason per node group - one node has a taint (a "keep out" mark, lesson 17.13) it does not tolerate, two have Insufficient cpu. The preemption: part says the scheduler also could not make room by removing lower-priority pods (lesson 17.18).
And the opposite: a pod with no requests reserves nothing, so the scheduler will happily stack twenty of them on a node that then runs out of real memory. "Insufficient cpu" means requested, never used.
(Total limits may be over 100 percent, i.e., overcommitted.) is kubectl telling you the other half: limits are not reserved at all. The sum of limits on a node routinely exceeds 100% - that is overcommit, and it is normal. It only hurts when enough pods actually use their limits at the same time.
Compare with what is really used right now. kubectl top node shows live usage; it needs metrics-server, a small add-on that asks every kubelet for CPU and memory usage (sampled every 15s) and serves it to kubectl:
$ kubectl top node
NAME CPU(cores) CPU(%) MEMORY(bytes) MEMORY(%)
cp-1 173m 9% 1388Mi 36%
worker-1 67m 3% 656Mi 17%
worker-2 114m 6% 1019Mi 27%
top percentages are against allocatable, not against requests. The gap between "Requests 22%" and "CPU 3%" is money: requests you pay for (nodes) and do not use. Right-sizing requests is one of the first things a platform team does.
What the kubelet does with the numbers: cgroups
The kubelet turns each container's resources into cgroup v2 settings (the same files you read under /sys/fs/cgroup in 5.11 and 11.10). You can read them from inside the container: kubectl exec <pod> -n <ns> -- <command> runs a command in the container (15.42), and there - thanks to a private cgroup namespace (containerd's default: the container sees only its own branch of the cgroup tree) - /sys/fs/cgroup is the container's own cgroup:
# shop/web with the resources above (an illustration; the next mission sets them for real)
kubectl exec web-5d8f-2kq9x -n shop -- cat /sys/fs/cgroup/cpu.max
50000 100000
kubectl exec web-5d8f-2kq9x -n shop -- cat /sys/fs/cgroup/cpu.weight
4
kubectl exec web-5d8f-2kq9x -n shop -- cat /sys/fs/cgroup/memory.max
268435456
cpu.max= quota period in microseconds.50000 100000= 50 ms of CPU time per 100 ms CFS period = limits.cpu 500m. No CPU limit printsmax 100000.cpu.weightcomes from requests.cpu (100m -> weight 4; 1 core -> 39; the scale is 1-10000). It only matters when the node's CPU is contended: then each cgroup gets CPU in proportion to its weight. Uncontended, a pod with a 100m request can burst to its limit (or the whole node, with no limit).memory.max= limits.memory in bytes (256Mi = 268435456). No limit printsmax. Memory requests do not appear in the cgroup at all - they are used by the scheduler and by the kubelet's eviction ranking (next lesson).
This is the same cgroup machinery as Ch 2's (lesson 2.26) MemoryMax= and CPUQuota= on a systemd unit - systemd-run -p MemoryMax=200M killed memhog with 137 the same way a container limit does. Kubernetes just writes the files for you.
Units, and the mistakes they cause
cpu: 1 = 1000m = one core (vCPU). 0.5 = 500m. "100m" = a tenth of a core.
memory: Mi/Gi = powers of 1024. M/G = powers of 1000. 128Mi = 134217728 bytes.
memory: 128m # 0.128 BYTES - a lowercase m is milli, not mega
memory: 1G # 1000000000 bytes, ~7% less than 1Gi
cpu: 1000 # one thousand cores
The apiserver accepts all three - they are valid quantities - and the pod either never schedules (1000 cores) or is killed immediately (0.128 bytes). Always use Mi/Gi for memory and m for CPU.
Two validation errors you will hit when editing by hand:
The Pod "web" is invalid: spec.containers[0].resources.requests: Invalid value: "2": must be less than or equal to cpu limit of 1
The Deployment "web" is invalid: spec.template.spec.containers[0].resources.limits[memory]: Invalid value: "512mb": must be a valid quantity
A request may never exceed its limit. And if you set only a limit, the apiserver copies it into the request - which quietly makes that pod Guaranteed (next lesson) and reserves the full limit on the node.
Setting them
# shop/web with the resources above (an illustration; the next mission sets them for real)
kubectl set resources deploy/web -n shop -c web --requests=cpu=100m,memory=128Mi --limits=cpu=500m,memory=256Mi
deployment.apps/web resource requirements updated
kubectl set resources edits the resources in a workload's pod template: deploy/web = the Deployment called web; -n shop = in namespace shop; -c web = only the container named web; --requests= / --limits= = comma-separated resource=amount pairs. The alternative is editing the YAML (kubectl edit, 15.40) - same result.
Changing resources changes the pod template, so it is a rollout: new pods, new cgroups. (In-place pod resize - beta in 1.33, GA in 1.35 - resizes a running container's cgroup without a restart via the pod's resize subresource, but changing a Deployment's template still rolls new pods.)
Pod-level totals: the scheduler adds up the containers. Init containers run one at a time, so a pod's effective request is max(largest init container, sum of app containers) - plus sidecars (init containers with restartPolicy: Always), which run alongside and are added to the sum.
What you can now do
- Explain the difference: requests -> placement and fair share. limits -> hard cgroup ceilings.
- "Insufficient cpu/memory" is always about requests against allocatable.
- A pod without requests is invisible to the scheduler's arithmetic. That is not "free" - it is a pod the scheduler will overpack and the kubelet will evict first.
- Read a node's ledger (
describe node), its live usage (top node) and a container's cgroup files, and set resources withkubectl set resources.