OnCallReady

Lesson 17.6 · Kubernetes: Scheduling, Health & Security · 19 min read

CPU throttling vs memory OOMKill: two failure modes, two fixes

In plain words

Imagine two kinds of school rules. Rule one: "you may use the playground for 10 minutes out of every hour". If your class is fast and uses its 10 minutes in the first bit of the hour, you just wait inside until the next hour starts. Nobody gets hurt; you're just slowed down. Rule two: "your locker holds exactly ten books". Try to put in an eleventh and the locker bursts and everything falls out.

CPU is the playground: over the limit, the container is throttled, paused for the rest of each 100 ms period, so it gets slow but never restarts. Memory is the locker: over the limit, the kernel OOM-kills it, exit 137, and it restarts. cpu.stat counts the pauses (nr_throttled); OOMKilled in Last State shows the burst.

Compressible vs incompressible

The problem. Two tickets look alike - "the service is slow" and "the service keeps restarting" - and both get blamed on "Kubernetes". One is the CPU limit freezing the app 100 ms at a time, the other is the memory limit killing it. This lesson teaches you to tell them apart and fix each one.

What you need to know already: requests, limits, cpu.max and the CFS period (17.1), QoS and eviction (17.3), signals and exit codes 128+N (3.6), cgroup OOM (5.11), the JVM's heap vs the rest of its memory (5.13).

CPU is compressible: if you want more than you are allowed, you simply get less and run slower. Memory is incompressible: a page you need either exists or it does not, and the kernel cannot "give you less" of it. Everything in this lesson follows from that:

over the CPU limitover the memory limit
what happensthrottled: the cgroup is paused until the next periodOOM-killed: the kernel kills a process in the cgroup
visible aslatency, timeouts, slow startuprestarts, exit 137, OOMKilled
restarts?neveryes, every time
kubectl toppinned at the limitclimbs, then drops to zero
where to lookcpu.stat inside the containerdescribe pod Last State, dmesg/journal on the node
fixraise or remove the CPU limit, or make the code cheaperraise the limit or make the app use less (heap!)

How throttling works: 100 ms at a time

A CPU limit is a CFS quota per period. The period is 100 ms. limits.cpu: 500m = 50 ms of CPU time per 100 ms, summed over all threads.

# an illustration: fx/quotes and pay/ledger are the throttling mission's pods
kubectl exec quotes-6d9f-2mxkq -n fx -- cat /sys/fs/cgroup/cpu.max
100000 100000

(1 core: 100000 us per 100000 us period.) Now the part that surprises everyone: a service with 8 busy threads and a 1-core limit burns its whole 100 ms quota in the first 12.5 ms of the period and is then frozen for the remaining 87.5 ms. Every request that arrived in that window waits. Average CPU over a minute looks calm - 40% of the limit - and p99 latency (the time within which 99% of requests finish, 0.2) has 80-100 ms spikes that repeat every period.

"A pod has limits.cpu: 1, average CPU is 40%, and p99 latency spikes every 100 ms" - that is throttling. Averages hide bursts; the quota is enforced per 100 ms, not per minute.

The kernel counts it for you, per cgroup:

# an illustration: fx/quotes and pay/ledger are the throttling mission's pods
kubectl exec quotes-6d9f-2mxkq -n fx -- cat /sys/fs/cgroup/cpu.stat
usage_usec 14400000
user_usec 11808000
system_usec 2592000
nr_periods 360
nr_throttled 108
throttled_usec 3800016
nr_bursts 0
burst_usec 0

108 / 360 = 30% of periods throttled. Anything above a few percent on a latency-sensitive service is worth fixing. Read it twice, 30 seconds apart, and subtract - the counters are cumulative since the container started (a restart is a new cgroup and they reset to 0).

Later (Ch 27): monitoring systems collect these same two counters for every container (container_cpu_cfs_periods_total, container_cpu_cfs_throttled_periods_total) and graph the throttled ratio for you.

Meanwhile kubectl top tells you almost nothing:

# an illustration: fx/quotes and pay/ledger are the throttling mission's pods
kubectl top pod -n fx
NAME                CPU(cores)   MEMORY(bytes)
quotes-6d9f-2mxkq   392m         24Mi

392m of a 1-core limit looks like lots of headroom. It is an average over the metrics-server window. That is why throttling is "the most under-diagnosed performance problem in Kubernetes".

The case against CPU limits

Many teams now run latency-sensitive services with CPU requests but no CPU limits. The argument:

The counter-arguments (and why banks often keep limits): noisy neighbours (one greedy pod slowing everyone on a shared node), predictable capacity planning, namespace ResourceQuotas (budgets per namespace, lesson 17.9) that require limits.cpu, and JVMs sizing their thread pools from the CPU limit (with no limit the JVM sees every core of the node). The defensible position: memory limits always; CPU limits deliberately - off for latency-critical services that are properly requested, on for batch and for multi-tenant namespaces.

OOMKilled: what it looks like

# an illustration: fx/quotes and pay/ledger are the throttling mission's pods
kubectl get pods -n pay
NAME                      READY   STATUS             RESTARTS      AGE
ledger-7b9c-kq2x8         0/1     CrashLoopBackOff   4 (41s ago)   6m
# an illustration: fx/quotes and pay/ledger are the throttling mission's pods
kubectl describe pod ledger-7b9c-kq2x8 -n pay
    State:          Waiting
      Reason:       CrashLoopBackOff
    Last State:     Terminated
      Reason:       OOMKilled
      Exit Code:    137
      Started:      Thu, 24 Sep 2026 10:14:02 +0000
      Finished:     Thu, 24 Sep 2026 10:15:11 +0000
    Ready:          False
    Restart Count:  4
    Limits:
      memory:  512Mi

The diagnosis path for "pod restarts every few hours with exit 137": describe (OOMKilled?) -> kubectl top pod over time (steady climb = leak or unbounded cache; sudden jump = a big request/batch) -> compare with the limit -> for a JVM, compare heap settings with the limit.

The JVM trap

A JVM (Java Virtual Machine - the program that runs Java applications, 5.13) uses memory as heap + everything else. The heap is where the app's data objects live, capped by -Xmx. "Everything else": metaspace (descriptions of the app's code), thread stacks (1 MiB per thread by default), code cache (compiled code), GC structures (bookkeeping of the garbage collector, which frees unused heap), direct buffers. It is typically 100-300 MiB for a Spring Boot service (a common Java web framework).

resources:
  limits:
    memory: 512Mi
env:
- name: JAVA_TOOL_OPTIONS
  value: "-Xmx512m"          # heap alone may grow to the whole limit

JAVA_TOOL_OPTIONS is an environment variable every JVM reads at startup and adds to its command line - the easy way to pass JVM flags to a container you did not build.

The heap grows towards -Xmx (the JVM is lazy about collecting while it may still grow), non-heap sits on top, the cgroup total crosses 512Mi, and the kernel kills the process. Not a Java OutOfMemoryError - an OOM kill, no stack trace, exit 137.

Modern JVMs (11+) are container-aware: they read memory.max from the cgroup. Without -Xmx, the default max heap is 25% of the container limit (-XX:MaxRAMPercentage=25) - safe, but often too small, and then you get the other failure: java.lang.OutOfMemoryError: Java heap space, which with -XX:+ExitOnOutOfMemoryError is a plain exit code 3 (not 137). The usual container setting:

env:
- name: JAVA_TOOL_OPTIONS
  value: "-XX:MaxRAMPercentage=75"   # heap = 75% of the limit, 25% left for the rest

You can see what the JVM picked up in the first log line:

Picked up JAVA_TOOL_OPTIONS: -XX:MaxRAMPercentage=75

"How do you size a container limit for a JVM?" -> measure the heap the app actually needs under load, add the non-heap overhead (measure it: NMT, jcmd VM.native_memory, or container memory minus heap), set the limit to their sum plus a margin, and express the heap as a percentage of the limit so the two can never disagree. (NMT = Native Memory Tracking, a JVM feature that reports the non-heap memory by category.)

Later (Ch 20): the JVM chapter measures heap, metaspace and thread stacks with the JDK's own tools.

Telling them apart in practice

restarts climbing, 137, OOMKilled   -> memory: limit vs real usage (heap?)
restarts climbing, 143 or 137 "Error", "Liveness probe failed" events -> probes (17.20)
no restarts, slow, timeouts, top near the limit or bursty -> CPU throttling
pod Failed "Evicted", node MemoryPressure events -> node-level memory, QoS

What you can now do

Why it helps

These two failures produce the tickets you'll see most. "The service is slow, p99 spikes, but CPU is only at 40%": throttling, invisible in averages, visible in cpu.stat or container_cpu_cfs_throttled_periods_total. "The pod restarts every few hours with 137": an OOM kill, usually a JVM whose -Xmx equals the container limit, so heap plus non-heap crosses it.

Knowing the mechanism lets you pick the right fix instead of "give it more of everything": raise or remove the CPU limit for the first, fix heap sizing with MaxRAMPercentage for the second. The JVM trap is also a near-certain interview question for anyone working with Java platforms, and the CPU-limits debate comes up in every platform team's resource policy.

FAQ

How can a pod be throttled at 40% average CPU?

The quota is enforced per 100 ms period, summed over all threads. A service with 8 busy threads and a 1-core limit uses its whole 100 ms of CPU time in the first 12.5 ms of each period and then waits 87.5 ms. Averaged over a minute that looks calm; requests arriving in the frozen window see 80-100 ms of extra latency. cpu.stat shows it as a high nr_throttled / nr_periods ratio.

Why is there no Java OutOfMemoryError when my JVM gets OOMKilled?

Because the kernel killed it, not the JVM. An OOMKill happens when the whole container, heap plus metaspace, threads, code cache and direct buffers, crosses the cgroup's memory.max. The kernel sends SIGKILL, so there's no stack trace, just exit 137. A Java OutOfMemoryError: Java heap space is the other failure: the heap hit -Xmx while the container was still under its limit.

Should we remove CPU limits?

For latency-critical services with properly set requests, many teams do: the request already guarantees a fair share under contention, and a limit only forbids using idle CPU while causing throttling. Keep them for batch work, noisy multi-tenant namespaces, quotas that require them, and JVMs that size thread pools from the CPU limit. Memory limits, always.

What JVM setting should I use in a container?

Let the heap follow the container limit: -XX:MaxRAMPercentage=75 (for example through JAVA_TOOL_OPTIONS) and no -Xmx, so heap is 75% of the limit and the rest is left for non-heap. Modern JVMs read the cgroup limit; without a setting the default heap is only 25% of it. Measure non-heap under load and adjust the percentage.

Why do the cpu.stat counters reset?

They're cumulative per cgroup, and a restarted container gets a new cgroup, so they start at zero again. To measure the current throttling rate, read them twice some seconds apart and subtract; monitoring systems do the same subtraction for you over the same two counters.

In an interview Mid

What is the difference between hitting a CPU limit and hitting a memory limit?

CPU is compressible, memory is not.

For a JVM, the classic 137 is a heap allowed to grow to the limit with no room for metaspace, thread stacks and the rest: size the limit as heap + non-heap + margin, and express the heap as a percentage (-XX:MaxRAMPercentage=75) so they cannot disagree.

Also asked: A Java service restarts every few hours with exit code 137. How do you diagnose it? · Should you set CPU limits on latency-sensitive services? · Why does kubectl top not show CPU throttling?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.