Cachingrediscachinghot-keythundering-herdpostgreshard

80% of reads hit one post ID. Redis pins at 100% and every key's p99 goes 1ms → 300ms. Then the TTL expires.

01Symptom

A single Redis instance sits between 30 app servers and Postgres for 'get post by ID'. Traffic is normally evenly spread across millions of posts and Redis idles. A celebrity post goes viral; within minutes 80% of all reads target that one ID. Redis CPU hits 100% and p99 latency jumps from 1ms to 300ms — for every key, not just the viral one. A few minutes later the hot entry expires at peak traffic (flat 60s TTL) and Postgres CPU spikes to 100% in the same instant, causing a brief full outage.

02Constraints

  • One Redis instance, not clustered; 4 vCPUs; Redis is single-threaded for command execution (I/O threads, when enabled, only move bytes across sockets — every command still executes on the one main thread)
  • Normal peak: 50,000 reads/sec, spread across ~10M distinct post IDs
  • Viral spike: 200,000 reads/sec total, of which 160,000/sec for a single post ID
  • Cache TTL is a flat 60s for all post entries — set by the same writer, so every expiry is synchronized with its own writes
  • Postgres can sustain ~2,000 reads/sec comfortably before latency degrades
  • 30 app servers behind a load balancer; no internal coordination between them (no shared singleflight, no shared backfill lock)
  • You cannot predict in advance which post will go viral
  • The viral post exists in Redis (it was written normally and served with the rest until its reads exploded)

03Evidence

  • Redis command-rate graph during the spike: ~160K of the ~200K ops/sec are GET for the same key (same post ID); the other 40K are spread across the rest of the keyspace
  • Redis CPU pinned at 100% the whole window, while memory is fine — no evictions, no OOM, no maxmemory activity
  • Redis p99 is uniform across ALL keys during the spike: the slow requests are not the viral key — GETs for cold, unrelated, even nonexistent keys all share the same ~300ms tail
  • The instant before the hot key's 60s TTL expires, Postgres CPU is steady (<10%); the second after it expires, Postgres CPU jumps to 100% and every app server logs the same SELECT for the same post ID starting within milliseconds of each other
  • APM shows thousands of identical in-flight queries for the same post row, all started in the same second, no request coalescing anywhere in the stack

→The question

Reads for one key at 160K/sec slow down requests for every other key in the same Redis — why, when Redis is otherwise 'fast enough' for the volume? And what exactly breaks in the seconds after that key's TTL hits zero?

04Your prediction

01Why do unrelated keys get slow?
02What happens in the seconds after the hot key's TTL expires?
03Is the fix about Redis, Postgres, or the request path between them?
04Would adding more Redis replicas fix this? Would sharding?
05The viral post gets 160K LIKEs/min. How is a write-heavy hot key different from this read-heavy case?
0 / 600 chars