OnCallReady

Lesson 20.12 · JVM Internals · 10 min read

Garbage collection: pauses, minor vs full, and the collectors

In plain words

Imagine a playroom where kids keep pulling out toys. Every so often a parent says "freeze!" so they can collect the toys nobody is playing with. A quick tidy of the toy box by the door takes a second. A full clean-up of the whole room takes much longer, and nobody may move until it is done. If the room is full of toys everybody is still using, the parent keeps shouting "freeze!" and cleans almost nothing.

The garbage collector is that parent, and "freeze" is a stop-the-world pause at a safepoint. Young GCs are the quick tidy, full GCs are the long clean-up. G1, the default, tries to keep each freeze under 200 ms and does much of its old-generation work while the kids keep playing.

What a pause is

The problem. "The service freezes for a second every few minutes" and "CPU is at 100% but traffic is normal" are often the garbage collector, not the code. To tell, you need to know what the collector does and what its work looks like from outside.

What you need to know already: the heap, young and old generation, eden and survivors (20.1), ergonomics picking a collector (20.2), percentiles and latency (0.2).

Stop-the-world = a pause in which every application thread is frozen so the collector can work safely. The collectors differ mostly in how long and how often those pauses are.

To move objects safely the collector must stop the application threads at a safepoint - "stop the world". A pause is the time from the last application thread stopping to the first one resuming. During a pause no request makes progress: a 400 ms pause is 400 ms added to every in-flight request, and to your p99.

Collectors differ in how much work they do inside the pause and how much they do concurrently, while the application runs.

Minor, mixed, full

young (minor) GC   collect eden + survivors         frequent, 2-20 ms      normal
concurrent cycle   mark the old gen while running   no pause (tiny ones)   normal
mixed GC (G1)      young + some old regions         a bit longer            normal
full GC            everything, compacting, STW      100 ms - many seconds   a symptom

A young GC every few seconds under load is healthy. Its cost is proportional to survivors, not to allocation. Frequent Full GCs mean the old generation is full: a leak, a heap too small for the live set, or a burst that promoted too much.

G1, the default

G1 ("garbage first") divides the heap into regions and tracks how much garbage each holds. It aims for a pause target (-XX:MaxGCPauseMillis=200 by default) by choosing how much to collect each time.

1. young GCs            copy live eden objects to survivor or old regions
2. old occupancy > IHOP (45% by default, adaptive)
   -> Pause Young (Concurrent Start)   a young GC that also starts marking
   -> Concurrent Mark Cycle            marks live objects while the app runs
   -> Pause Remark, Pause Cleanup      two short pauses to finish marking
3. Pause Young (Prepare Mixed), then several Pause Young (Mixed)
                        young + the old regions with the most garbage
4. if marking or mixed collections cannot keep up:
   -> Pause Full (G1 Compaction Pause)   the fallback, all threads stopped

Tuning G1 in 2026 mostly means not tuning it: give it enough heap, set a pause target if the default does not suit you, leave the rest.

The others

Serial     -XX:+UseSerialGC     one thread, STW. Picked automatically below 2 CPUs / 1792 MB.
Parallel   -XX:+UseParallelGC   many threads, STW, best throughput. Batch jobs.
G1         -XX:+UseG1GC         default. Region-based, pause target, mostly concurrent old gen.
ZGC        -XX:+UseZGC          concurrent almost everything: sub-millisecond pauses at any heap size.
                                JDK 21: generational mode needs -XX:+ZGenerational; default from JDK 23.
Shenandoah -XX:+UseShenandoahGC concurrent compaction, low pauses. Not in every JDK build.

ZGC and Shenandoah trade some throughput and memory headroom for pauses that do not grow with the heap. Worth it for latency-critical services with big heaps; not a fix for a leak or an undersized heap. A low-pause collector with a leak still runs out of memory - it just gets there with nicer pause graphs.

How GC shows up to you

The flags you add to every service

-XX:MaxRAMPercentage=75
-XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=/var/tmp/
-XX:+ExitOnOutOfMemoryError
-Xlog:gc*:file=/var/log/app/gc.log:time,uptime:filecount=5,filesize=10M

Common misreadings

Why it helps

When p99 latency spikes every few minutes for no obvious reason, GC pauses are the first suspect, and you can line them up with the GC log's pause times. When a Java service sits at 100% CPU and serves nothing, it is very often the collector thrashing in full GC after full GC, not application code. Knowing that saves hours of profiling the wrong thing.

It also helps in design discussions. Someone will propose switching to ZGC to fix OutOfMemoryErrors; you know a low-pause collector does not fix a leak or an undersized heap. And you will add the standard flags to every service template: HeapDumpOnOutOfMemoryError, ExitOnOutOfMemoryError and rotating GC logs, so the next incident leaves evidence behind.

FAQ

Are frequent young GCs a problem?

No. A young GC every few seconds under load is normal and healthy. Its cost depends on how many objects survive, not how many were allocated, so each one usually takes a few milliseconds. What you look at is the total time spent in GC as a fraction of wall time, the longest pauses, and whether any full GCs happen. Thousands of young GCs a day with low total time are fine.

What is the difference between a mixed GC and a full GC in G1?

A mixed GC is a normal pause that collects the young generation plus some old regions with the most garbage, after a concurrent marking cycle has identified them. It is part of G1 working as designed. A full GC, Pause Full (G1 Compaction Pause), is the fallback when marking or mixed collections cannot keep up: it stops everything and compacts the whole heap. In G1 a full GC is a symptom worth investigating.

Should I switch to ZGC?

Only for the right problem. ZGC keeps pauses below a millisecond regardless of heap size, which is valuable for latency-critical services with large heaps. It costs some throughput and needs more memory headroom. It does nothing for a leak or a heap that is too small for its live set; it just gets to the OutOfMemoryError with nicer pause graphs. In JDK 21 you add -XX:+ZGenerational; from JDK 23 generational is the default.

Why would I want the JVM to exit on OutOfMemoryError?

Because after an OOM the JVM can be half broken: some threads died with the error, pools may be in odd states, and the service may still answer health checks. -XX:+ExitOnOutOfMemoryError makes it exit with code 3 on the first OOM so systemd or Kubernetes restarts it cleanly. Pair it with HeapDumpOnOutOfMemoryError so the evidence is written first. It is not manageable, so it needs a restart to set.

Is calling System.gc() a good way to free memory?

No. It triggers a full GC, which stops the application and usually frees nothing that the normal collections would not have freed anyway. In production code it is almost always a bug, often in a library. If you see System.gc() as a cause in the GC log and cannot remove it, -XX:+DisableExplicitGC turns those calls into no-ops. It also does not return memory to the container in a way that helps with limits.

In an interview Mid

What is the difference between a minor GC and a full GC, and when is GC a problem?

Most objects die young, so the heap has a young generation (eden + survivors) and an old generation.

GC is a problem when it shows up from outside: latency spikes that line up with pauses (a 400 ms pause adds 400 ms to every in-flight request), or 100% CPU with throughput near zero - Full GC after Full GC freeing almost nothing. A low-pause collector like ZGC does not fix a leak.

Also asked: What is a stop-the-world pause? · Compare the Serial, Parallel, G1 and ZGC collectors at a high level. · Which GC-related JVM flags would you put on every production service?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.