Compressible vs incompressible
The problem. Two tickets look alike - "the service is slow" and "the service keeps restarting" - and both get blamed on "Kubernetes". One is the CPU limit freezing the app 100 ms at a time, the other is the memory limit killing it. This lesson teaches you to tell them apart and fix each one.
What you need to know already: requests, limits, cpu.max and the CFS period (17.1), QoS and eviction (17.3), signals and exit codes 128+N (3.6), cgroup OOM (5.11), the JVM's heap vs the rest of its memory (5.13).
CPU is compressible: if you want more than you are allowed, you simply get less and run slower. Memory is incompressible: a page you need either exists or it does not, and the kernel cannot "give you less" of it. Everything in this lesson follows from that:
| over the CPU limit | over the memory limit | |
|---|---|---|
| what happens | throttled: the cgroup is paused until the next period | OOM-killed: the kernel kills a process in the cgroup |
| visible as | latency, timeouts, slow startup | restarts, exit 137, OOMKilled |
| restarts? | never | yes, every time |
kubectl top | pinned at the limit | climbs, then drops to zero |
| where to look | cpu.stat inside the container | describe pod Last State, dmesg/journal on the node |
| fix | raise or remove the CPU limit, or make the code cheaper | raise the limit or make the app use less (heap!) |
How throttling works: 100 ms at a time
A CPU limit is a CFS quota per period. The period is 100 ms. limits.cpu: 500m = 50 ms of CPU time per 100 ms, summed over all threads.
# an illustration: fx/quotes and pay/ledger are the throttling mission's pods
kubectl exec quotes-6d9f-2mxkq -n fx -- cat /sys/fs/cgroup/cpu.max
100000 100000
(1 core: 100000 us per 100000 us period.) Now the part that surprises everyone: a service with 8 busy threads and a 1-core limit burns its whole 100 ms quota in the first 12.5 ms of the period and is then frozen for the remaining 87.5 ms. Every request that arrived in that window waits. Average CPU over a minute looks calm - 40% of the limit - and p99 latency (the time within which 99% of requests finish, 0.2) has 80-100 ms spikes that repeat every period.
"A pod has
limits.cpu: 1, average CPU is 40%, and p99 latency spikes every 100 ms" - that is throttling. Averages hide bursts; the quota is enforced per 100 ms, not per minute.
The kernel counts it for you, per cgroup:
# an illustration: fx/quotes and pay/ledger are the throttling mission's pods
kubectl exec quotes-6d9f-2mxkq -n fx -- cat /sys/fs/cgroup/cpu.stat
usage_usec 14400000
user_usec 11808000
system_usec 2592000
nr_periods 360
nr_throttled 108
throttled_usec 3800016
nr_bursts 0
burst_usec 0
nr_periods- periods in which the cgroup had runnable work (with a quota set)nr_throttled- periods in which it hit the quota and was pausedthrottled_usec- total time spent paused
108 / 360 = 30% of periods throttled. Anything above a few percent on a latency-sensitive service is worth fixing. Read it twice, 30 seconds apart, and subtract - the counters are cumulative since the container started (a restart is a new cgroup and they reset to 0).
Later (Ch 27): monitoring systems collect these same two counters for every container (
container_cpu_cfs_periods_total,container_cpu_cfs_throttled_periods_total) and graph the throttled ratio for you.
Meanwhile kubectl top tells you almost nothing:
# an illustration: fx/quotes and pay/ledger are the throttling mission's pods
kubectl top pod -n fx
NAME CPU(cores) MEMORY(bytes)
quotes-6d9f-2mxkq 392m 24Mi
392m of a 1-core limit looks like lots of headroom. It is an average over the metrics-server window. That is why throttling is "the most under-diagnosed performance problem in Kubernetes".
The case against CPU limits
Many teams now run latency-sensitive services with CPU requests but no CPU limits. The argument:
- The request already guarantees the pod its share under contention (cpu.weight). The limit adds nothing for fairness - it only forbids using idle CPU.
- Throttling is invisible in averages and hurts tail latency, the thing users feel.
- Idle CPU is worth zero; letting a pod burst into it is free.
The counter-arguments (and why banks often keep limits): noisy neighbours (one greedy pod slowing everyone on a shared node), predictable capacity planning, namespace ResourceQuotas (budgets per namespace, lesson 17.9) that require limits.cpu, and JVMs sizing their thread pools from the CPU limit (with no limit the JVM sees every core of the node). The defensible position: memory limits always; CPU limits deliberately - off for latency-critical services that are properly requested, on for batch and for multi-tenant namespaces.
OOMKilled: what it looks like
# an illustration: fx/quotes and pay/ledger are the throttling mission's pods
kubectl get pods -n pay
NAME READY STATUS RESTARTS AGE
ledger-7b9c-kq2x8 0/1 CrashLoopBackOff 4 (41s ago) 6m
# an illustration: fx/quotes and pay/ledger are the throttling mission's pods
kubectl describe pod ledger-7b9c-kq2x8 -n pay
State: Waiting
Reason: CrashLoopBackOff
Last State: Terminated
Reason: OOMKilled
Exit Code: 137
Started: Thu, 24 Sep 2026 10:14:02 +0000
Finished: Thu, 24 Sep 2026 10:15:11 +0000
Ready: False
Restart Count: 4
Limits:
memory: 512Mi
- 137 = 128 + 9 = killed by SIGKILL (Ch 3's signal arithmetic, 3.6).
- CrashLoopBackOff = the container keeps dying and the kubelet waits longer and longer (10s, 20s, 40s... up to 5 min) before each restart (15.14).
Reason: OOMKilledis how you know it was the cgroup OOM killer and not akubectl delete, a liveness kill (a failed health check, lesson 17.20) or a node eviction (those would beErrorwith 137 or 143, or an Evicted pod).- The evidence is in Last State of the current container. The new container's own
memory.eventsstarts from zero - a restart is a new cgroup. - The kernel logs it on the node:
Memory cgroup out of memory: Killed process 1047 (java) total-vm:...(journalctl -kon the node - the kernel log, 5.8).
The diagnosis path for "pod restarts every few hours with exit 137": describe (OOMKilled?) -> kubectl top pod over time (steady climb = leak or unbounded cache; sudden jump = a big request/batch) -> compare with the limit -> for a JVM, compare heap settings with the limit.
The JVM trap
A JVM (Java Virtual Machine - the program that runs Java applications, 5.13) uses memory as heap + everything else. The heap is where the app's data objects live, capped by -Xmx. "Everything else": metaspace (descriptions of the app's code), thread stacks (1 MiB per thread by default), code cache (compiled code), GC structures (bookkeeping of the garbage collector, which frees unused heap), direct buffers. It is typically 100-300 MiB for a Spring Boot service (a common Java web framework).
resources:
limits:
memory: 512Mi
env:
- name: JAVA_TOOL_OPTIONS
value: "-Xmx512m" # heap alone may grow to the whole limit
JAVA_TOOL_OPTIONS is an environment variable every JVM reads at startup and adds to its command line - the easy way to pass JVM flags to a container you did not build.
The heap grows towards -Xmx (the JVM is lazy about collecting while it may still grow), non-heap sits on top, the cgroup total crosses 512Mi, and the kernel kills the process. Not a Java OutOfMemoryError - an OOM kill, no stack trace, exit 137.
Modern JVMs (11+) are container-aware: they read memory.max from the cgroup. Without -Xmx, the default max heap is 25% of the container limit (-XX:MaxRAMPercentage=25) - safe, but often too small, and then you get the other failure: java.lang.OutOfMemoryError: Java heap space, which with -XX:+ExitOnOutOfMemoryError is a plain exit code 3 (not 137). The usual container setting:
env:
- name: JAVA_TOOL_OPTIONS
value: "-XX:MaxRAMPercentage=75" # heap = 75% of the limit, 25% left for the rest
You can see what the JVM picked up in the first log line:
Picked up JAVA_TOOL_OPTIONS: -XX:MaxRAMPercentage=75
"How do you size a container limit for a JVM?" -> measure the heap the app actually needs under load, add the non-heap overhead (measure it: NMT, jcmd VM.native_memory, or container memory minus heap), set the limit to their sum plus a margin, and express the heap as a percentage of the limit so the two can never disagree. (NMT = Native Memory Tracking, a JVM feature that reports the non-heap memory by category.)
Later (Ch 20): the JVM chapter measures heap, metaspace and thread stacks with the JDK's own tools.
Telling them apart in practice
restarts climbing, 137, OOMKilled -> memory: limit vs real usage (heap?)
restarts climbing, 143 or 137 "Error", "Liveness probe failed" events -> probes (17.20)
no restarts, slow, timeouts, top near the limit or bursty -> CPU throttling
pod Failed "Evicted", node MemoryPressure events -> node-level memory, QoS
What you can now do
- Measure throttling from
cpu.stat(two samples, subtract, divide). - Recognise an OOM kill in
describe podand explain 137. - Size a JVM container: heap as a percentage of the limit, room left for the rest.