What a pause is
The problem. "The service freezes for a second every few minutes" and "CPU is at 100% but traffic is normal" are often the garbage collector, not the code. To tell, you need to know what the collector does and what its work looks like from outside.
What you need to know already: the heap, young and old generation, eden and survivors (20.1), ergonomics picking a collector (20.2), percentiles and latency (0.2).
Stop-the-world = a pause in which every application thread is frozen so the collector can work safely. The collectors differ mostly in how long and how often those pauses are.
To move objects safely the collector must stop the application threads at a safepoint - "stop the world". A pause is the time from the last application thread stopping to the first one resuming. During a pause no request makes progress: a 400 ms pause is 400 ms added to every in-flight request, and to your p99.
Collectors differ in how much work they do inside the pause and how much they do concurrently, while the application runs.
Minor, mixed, full
young (minor) GC collect eden + survivors frequent, 2-20 ms normal
concurrent cycle mark the old gen while running no pause (tiny ones) normal
mixed GC (G1) young + some old regions a bit longer normal
full GC everything, compacting, STW 100 ms - many seconds a symptom
A young GC every few seconds under load is healthy. Its cost is proportional to survivors, not to allocation. Frequent Full GCs mean the old generation is full: a leak, a heap too small for the live set, or a burst that promoted too much.
G1, the default
G1 ("garbage first") divides the heap into regions and tracks how much garbage each holds. It aims for a pause target (-XX:MaxGCPauseMillis=200 by default) by choosing how much to collect each time.
1. young GCs copy live eden objects to survivor or old regions
2. old occupancy > IHOP (45% by default, adaptive)
-> Pause Young (Concurrent Start) a young GC that also starts marking
-> Concurrent Mark Cycle marks live objects while the app runs
-> Pause Remark, Pause Cleanup two short pauses to finish marking
3. Pause Young (Prepare Mixed), then several Pause Young (Mixed)
young + the old regions with the most garbage
4. if marking or mixed collections cannot keep up:
-> Pause Full (G1 Compaction Pause) the fallback, all threads stopped
Tuning G1 in 2026 mostly means not tuning it: give it enough heap, set a pause target if the default does not suit you, leave the rest.
The others
Serial -XX:+UseSerialGC one thread, STW. Picked automatically below 2 CPUs / 1792 MB.
Parallel -XX:+UseParallelGC many threads, STW, best throughput. Batch jobs.
G1 -XX:+UseG1GC default. Region-based, pause target, mostly concurrent old gen.
ZGC -XX:+UseZGC concurrent almost everything: sub-millisecond pauses at any heap size.
JDK 21: generational mode needs -XX:+ZGenerational; default from JDK 23.
Shenandoah -XX:+UseShenandoahGC concurrent compaction, low pauses. Not in every JDK build.
ZGC and Shenandoah trade some throughput and memory headroom for pauses that do not grow with the heap. Worth it for latency-critical services with big heaps; not a fix for a leak or an undersized heap. A low-pause collector with a leak still runs out of memory - it just gets there with nicer pause graphs.
How GC shows up to you
- Latency spikes that line up with GC pauses. Correlate p99 with
jvm_gc_pause_seconds_maxor the GC log. - High CPU that is the collector. A JVM near its heap limit runs Full GC after Full GC, each freeing almost nothing. CPU goes to 100%, throughput to zero, and the process is alive the whole time. "High CPU in a Java service is very often GC thrashing, not application work." Check GC time first.
- OutOfMemoryError: Java heap space after a long thrash - or never, if the collector keeps freeing just enough to limp on.
- "GC overhead limit exceeded" (Parallel): more than 98% of time in GC recovering less than 2% of the heap.
The flags you add to every service
-XX:MaxRAMPercentage=75
-XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=/var/tmp/
-XX:+ExitOnOutOfMemoryError
-Xlog:gc*:file=/var/log/app/gc.log:time,uptime:filecount=5,filesize=10M
HeapDumpOnOutOfMemoryErrorwrites the evidence at the moment of failure.HeapDumpPathas a directory givesjava_pid<PID>.hprofinside it.ExitOnOutOfMemoryErrormakes the JVM exit (code 3) on the first OOM instead of limping on in a half-broken state with some threads dead. Let systemd or Kubernetes restart it. It is not a manageable flag: it needs a restart to set.- GC logging is cheap enough to leave on in production. Rotating files keep it bounded.
Common misreadings
- "We had 3000 young GCs today." Normal. Look at the total time and the pauses.
- "Heap usage graph is a sawtooth going up to 95%." Normal: that is eden filling and emptying. Look at the troughs.
- "Switch to ZGC, it will fix the OOMs." It will not.
- "System.gc() will free memory." It triggers a Full GC. In production code it is almost always a bug;
-XX:+DisableExplicitGCexists for a reason.