Runtimejvmgarbage-collectionlatencyhard

CPU graph says 45%. p99 latency says otherwise.

01Symptom

A Java service on G1GC shows moderate average CPU (40-50%) — nothing alarming on the dashboard. But p99 latency has recurring spikes of 200-500ms every few seconds, invisible in the average, and users are noticing timeouts.

02Constraints

  • JVM service, G1GC with default tuning, fixed heap size
  • High allocation rate: request handlers create many short-lived DTOs and do string concatenation in the hot path
  • GC logs are enabled and available
  • Request rate is steady, no traffic spikes correlate with the latency spikes

03Evidence

  • GC logs show frequent young-generation collections, with pause durations growing over the service's uptime
  • Latency spike timestamps line up exactly with GC pause events in the GC log, to the millisecond
  • Allocation rate (bytes/sec, from JFR or GC logs) is very high relative to the service's actual throughput
  • Average CPU looks fine because GC pauses are infrequent relative to total time — they dominate the tail, not the mean

→The question

How do you confirm GC is the cause before changing any tuning flags, and what's actually driving the pause frequency?

04Your prediction

01What would you check first?
02What's actually driving the pause frequency?
0 / 600 chars