OnCallReady

Lesson 20.2 · JVM Internals · 17 min read

What the JVM decides at startup: container-aware ergonomics

In plain words

Imagine a new kid arriving at a sleepover. Nobody tells him how much of the pizza is his, so he looks around, counts the people and takes a careful quarter. If the room is small and only one other kid is there, he also decides to do things the simple way. He guesses from what he can see.

The JVM does the same at startup when you pass no flags. It reads how much memory and how many CPUs it can see, which in a container means the cgroup limits, and picks a heap of 25% of that and a collector. Under 2 CPUs or about 1.8 GB it picks Serial instead of G1. java -XX:+PrintFlagsFinal -version shows every decision, and systemd-run --scope -p MemoryMax=1G lets you watch it change.

Ergonomics

The problem. Start a Java service without memory flags and it picks a heap size by itself - usually a quarter of what it can see. In a small container that is wasteful; with the wrong flag it is fatal. You need to know what it picks and how to tell it otherwise.

What you need to know already: the heap and the rest of the footprint (20.1), cgroups and systemd-run --scope -p MemoryMax= (2.26, 5.11), JAVA_TOOL_OPTIONS and MaxRAMPercentage (17.6), systemd drop-ins (2.3).

Ergonomics = the JVM's word for the defaults it computes at startup. A flag is a JVM option on the command line: -Xmx768m (maximum heap 768 MB), -Xms (initial heap), and -XX:Name=value / -XX:+Name / -XX:-Name for the rest (+ switches a boolean on, - off).

When you do not pass a flag, the JVM chooses: which collector, how big the heap, how many GC threads. Those choices depend on the memory and CPUs it can see - and since JDK 10 (and 8u191) it reads them from the cgroup, not from the host. That is -XX:+UseContainerSupport, on by default.

The one command that shows every decision:

$ java -XX:+PrintFlagsFinal -version | grep -E 'MaxHeapSize|UseContainerSupport|MaxRAMPercentage'
   size_t MaxHeapSize                              = 1551892480                                {product} {ergonomic}
   double MaxRAMPercentage                         = 25.000000                                 {product} {default}
   size_t SoftMaxHeapSize                          = 1551892480                                {manageable} {ergonomic}
     bool UseContainerSupport                      = true                                      {product} {default}
openjdk version "21.0.8" 2025-07-15
OpenJDK Runtime Environment (build 21.0.8+9-Ubuntu-0ubuntu1~26.04)
OpenJDK 64-Bit Server VM (build 21.0.8+9-Ubuntu-0ubuntu1~26.04, mixed mode, sharing)

The columns: type, name, value, and where the value came from: {default} (compiled-in), {ergonomic} (computed at startup), {command line} (you set it). The version goes to stderr, which is why it survives the grep.

On the 6 GB host: MaxHeapSize = 25% of 5925 MB, rounded down to the heap alignment = 1480 MB. The default MaxRAMPercentage of 25 is deliberately conservative: the JVM assumes it shares the machine.

Inside a cgroup

systemd-run --scope puts a command in a transient cgroup with a limit - the closest thing on a VM to a container's memory limit:

$ sudo systemd-run --scope -p MemoryMax=1G java -XX:+PrintFlagsFinal -version | grep -E 'MaxHeapSize|UseSerialGC|UseG1GC'
Running as unit: run-r4e9e14955....scope; invocation ID: fdf574cd...
   size_t MaxHeapSize                              = 268435456                                 {product} {ergonomic}
     bool UseG1GC                                  = false                                     {product} {default}
     bool UseSerialGC                              = true                                      {product} {ergonomic}

Two things changed:

  1. MaxHeapSize is 256 MB - 25% of the 1 GB limit. A container with 1 GB and no flags gets a quarter of it as heap and wastes most of the rest.
  2. UseSerialGC is true. The JVM only picks G1 on a "server-class machine": at least 2 CPUs and at least 1792 MB (2 GB minus 256 MB) of memory. Under that, or with one CPU, you silently get the single-threaded Serial collector. Small pods are exactly where this bites.

The fix is to state what you want:

$ sudo systemd-run --scope -p MemoryMax=1G java -XX:MaxRAMPercentage=75 -XX:+UseG1GC -XX:+PrintFlagsFinal -version | grep -E ' MaxHeapSize|UseG1GC'
   size_t MaxHeapSize                              = 805306368                                 {product} {ergonomic}
     bool UseG1GC                                  = true                                      {product} {command line}

768 MB of heap in a 1 GB limit, 256 MB left for everything that is not heap.

What to check the JVM actually sees

$ sudo systemd-run --scope -p MemoryMax=1G java -XshowSettings:system -version
Operating System Metrics:
    Provider: cgroupv2
    Effective CPU Count: 2
    ...
    Memory Limit: 1.00G
    Memory Soft Limit: Unlimited
    Memory & Swap Limit: 1.00G

If Memory Limit says Unlimited inside a container, container support is off or the JVM is too old to read cgroup v2 (JDK 8 before 8u372, JDK 11 before 11.0.16). Then it sizes from the node's RAM - the classic "Java ate the node".

The flags that matter

-XX:MaxRAMPercentage=75           heap = 75% of the limit (use this, not -Xmx, in containers)
-XX:InitialRAMPercentage=50       start big to avoid early resizes (optional)
-Xmx768m                          fixed heap: wins over MaxRAMPercentage
-XX:+UseG1GC                      do not let a small limit pick Serial for you
-XX:MaxDirectMemorySize=256m      cap off-heap buffers (default is about -Xmx)
-Xss512k                          smaller thread stacks (many threads, shallow stacks)
-XX:ActiveProcessorCount=2        override what CPU count the JVM believes

MaxRAMPercentage vs -Xmx: with a percentage, the heap follows the limit when the platform team changes it; with -Xmx you have two numbers that must be changed together, and one day only one will be.

Why 75 and not 100: sizing

UseContainerSupport only sizes the heap. Everything from the previous lesson still needs room inside the same limit:

limit 1400 MB, MaxRAMPercentage=75
  heap                   1050 MB
  metaspace + classes     115 MB
  threads (48 x ~0.3)      15 MB
  code cache               48 MB
  GC, symbols, other       40 MB
                         -------
                         ~1270 MB   fits, 130 MB spare

Rules that survive contact with production:

A thread stack detail specific to this box

$ java -XX:+PrintFlagsFinal -version | grep ThreadStackSize
     intx ThreadStackSize                          = 2040                                      {pd product} {default}

On linux-aarch64 the default Java thread stack reservation is 2040 KB; on x86_64 it is 1024 KB. Reserved, not resident - only pages the thread actually touches count in RSS - but it is why "about 1 MB per thread" is only a rule of thumb.

Getting a flag into a running service

Four ways, least to most invasive:

Environment=JAVA_TOOL_OPTIONS=-XX:NativeMemoryTracking=summary    systemd drop-in
env: - name: JAVA_TOOL_OPTIONS  value: "-XX:MaxRAMPercentage=75"   Kubernetes
ExecStart= / ExecStart=/usr/bin/java -XX:... -jar app.jar          replace the command
jinfo -flag +HeapDumpOnOutOfMemoryError PID                        manageable flags only, live

Every JVM that sees JAVA_TOOL_OPTIONS prints it on stderr at startup - that line in the journal is your proof it was picked up:

java[17390]: Picked up JAVA_TOOL_OPTIONS: -XX:NativeMemoryTracking=summary

A typo in a flag is fatal, which is good - you find out at deploy time:

$ java -XX:MaxRamPercentage=75 -version
Unrecognized VM option 'MaxRamPercentage=75'
Error: Could not create the Java Virtual Machine.
Error: A fatal exception has occurred. Program will exit.

Flags are case-sensitive. -XX:+UseCGroupMemoryLimitForHeap from old blog posts was removed in JDK 10 and fails the same way.

Why it helps

Two production surprises come straight from ergonomics. A team moves a service to a 1 GB pod with no flags and gets a 256 MB heap: it runs out of heap under load while 700 MB of the limit sits unused. Another team gives a service limits.cpu: 1 and suddenly sees long pauses, because the JVM quietly switched to the single-threaded Serial collector.

When you review a Deployment's YAML, you can now ask the right questions: is MaxRAMPercentage set, is the collector explicit, does -Xmx fight the limit? And when someone says "Java ate the node", you know to check whether the JVM is old enough to read cgroup v2, with -XshowSettings:system.

Commands in this lesson

java systemd-run

FAQ

Why is the default heap only 25% of memory?

Because the default assumes the JVM shares the machine with other processes, as it did on servers before containers. On a 6 GB host that is about 1.5 GB of heap. In a container, the JVM is usually the only real process in the cgroup, so 25% wastes most of the limit. That is why you set -XX:MaxRAMPercentage=75 or so for containerised services instead of relying on the default.

Should I use -Xmx or MaxRAMPercentage in containers?

Prefer MaxRAMPercentage. With a percentage the heap follows the container limit: if the platform team raises the limit from 1 GiB to 2 GiB, the heap grows with it. With -Xmx you have two numbers that must be changed together, and sooner or later someone changes only the limit, or worse, raises -Xmx to match it. If both are set, -Xmx wins, so check for leftover -Xmx in the entrypoint or JAVA_TOOL_OPTIONS.

What is UseContainerSupport and do I need to turn it on?

It makes the JVM read memory and CPU limits from the cgroup instead of from the host. It is on by default since JDK 10 and 8u191, so you do not set it. What you do check is cgroup v2 support: JDK 8 before 8u372 and JDK 11 before 11.0.16 only understand cgroup v1, and on a modern node they report Memory Limit: Unlimited and size the heap from the node's RAM.

How does a flag get into a service without changing the image?

JAVA_TOOL_OPTIONS. Every JVM reads that environment variable and adds it to its flags, and prints Picked up JAVA_TOOL_OPTIONS: ... on stderr, which is your proof it took effect. On a systemd service you add it with a drop-in via systemctl edit; in Kubernetes as an env var in the pod spec. A few flags can be changed live with jinfo or jcmd VM.set_flag, but only those marked manageable.

What happens if I mistype a flag?

The JVM refuses to start: Unrecognized VM option 'MaxRamPercentage=75' followed by Could not create the Java Virtual Machine. Flags are case-sensitive, and some old ones from blog posts, like UseCGroupMemoryLimitForHeap, were removed and fail the same way. This is a good thing: you find out at deploy time rather than in production. In Kubernetes it shows up as a CrashLoopBackOff with that message in kubectl logs --previous.

In an interview Mid

How does the JVM decide its heap size in a container, and what would you set?

With container support (-XX:+UseContainerSupport, on by default since JDK 10) the JVM reads the memory and CPU limits from the cgroup and computes its defaults - ergonomics:

What I set: -XX:MaxRAMPercentage=75 (the heap follows the limit if someone changes it, unlike -Xmx), -XX:+UseG1GC explicitly, and never a heap equal to the limit - the rest of the footprint needs ~25%. Below ~512 MB, 50-60%.

How I verify: java -XX:+PrintFlagsFinal -version (the {ergonomic} / {command line} column says where each value came from), -XshowSettings:system (Memory Limit: Unlimited in a container = an old JDK sizing from the node), and Picked up JAVA_TOOL_OPTIONS in the log.

Also asked: Why prefer MaxRAMPercentage over -Xmx in a container? · How do you pass JVM flags to a service without changing its start command? · Why would a JVM in a small container end up on the Serial collector?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.