CPU graph says 45%. p99 latency says otherwise.
01Symptom
A Java service on G1GC shows moderate average CPU (40-50%) — nothing alarming on the dashboard. But p99 latency has recurring spikes of 200-500ms every few seconds, invisible in the average, and users are noticing timeouts.
02Constraints
- JVM service, G1GC with default tuning, fixed heap size
- High allocation rate: request handlers create many short-lived DTOs and do string concatenation in the hot path
- GC logs are enabled and available
- Request rate is steady, no traffic spikes correlate with the latency spikes
03Evidence
- GC logs show frequent young-generation collections, with pause durations growing over the service's uptime
- Latency spike timestamps line up exactly with GC pause events in the GC log, to the millisecond
- Allocation rate (bytes/sec, from JFR or GC logs) is very high relative to the service's actual throughput
- Average CPU looks fine because GC pauses are infrequent relative to total time — they dominate the tail, not the mean
→The question
How do you confirm GC is the cause before changing any tuning flags, and what's actually driving the pause frequency?
04Your prediction
0 / 600 chars