Cachingrediscachingpostgresmedium

98% cache hit rate → 40%, every 60 seconds, like clockwork

01Symptom

Your API caches expensive aggregation results in Redis with a flat 60-second TTL. Every 60 seconds, in a tight window, Postgres CPU spikes and p99 latency jumps from 20ms to 2-4 seconds — then it recovers, until the next cycle.

02Constraints

  • Redis cache, fixed 60s TTL set uniformly by a nightly batch job that warms the top 500 keys all within the same few seconds
  • No request coalescing — each cache miss triggers its own DB query
  • Traffic is steady, ~800 req/s, no deploys correlate with the spikes
  • Postgres is otherwise healthy: <20% CPU outside these windows

03Evidence

  • Redis hit rate graph is a sawtooth: 98% → 40% → 98%, period exactly 60s
  • Postgres query count graph spikes in the same windows, same queries repeated hundreds of times concurrently for the same cache keys
  • APM shows dozens of identical in-flight queries for the same aggregation, same parameters, all started within milliseconds of each other

→The question

What's causing the periodic spike, and why does it recur precisely every 60 seconds instead of randomly?

04Your prediction

01What's the root cause?
02Which fix addresses the root cause, not just the symptom?
0 / 600 chars