98% cache hit rate → 40%, every 60 seconds, like clockwork
01Symptom
Your API caches expensive aggregation results in Redis with a flat 60-second TTL. Every 60 seconds, in a tight window, Postgres CPU spikes and p99 latency jumps from 20ms to 2-4 seconds — then it recovers, until the next cycle.
02Constraints
- Redis cache, fixed 60s TTL set uniformly by a nightly batch job that warms the top 500 keys all within the same few seconds
- No request coalescing — each cache miss triggers its own DB query
- Traffic is steady, ~800 req/s, no deploys correlate with the spikes
- Postgres is otherwise healthy: <20% CPU outside these windows
03Evidence
- Redis hit rate graph is a sawtooth: 98% → 40% → 98%, period exactly 60s
- Postgres query count graph spikes in the same windows, same queries repeated hundreds of times concurrently for the same cache keys
- APM shows dozens of identical in-flight queries for the same aggregation, same parameters, all started within milliseconds of each other
→The question
What's causing the periodic spike, and why does it recur precisely every 60 seconds instead of randomly?
04Your prediction
0 / 600 chars