80% of reads hit one post ID. Redis pins at 100% and every key's p99 goes 1ms → 300ms. Then the TTL expires.
01Symptom
A single Redis instance sits between 30 app servers and Postgres for 'get post by ID'. Traffic is normally evenly spread across millions of posts and Redis idles. A celebrity post goes viral; within minutes 80% of all reads target that one ID. Redis CPU hits 100% and p99 latency jumps from 1ms to 300ms — for every key, not just the viral one. A few minutes later the hot entry expires at peak traffic (flat 60s TTL) and Postgres CPU spikes to 100% in the same instant, causing a brief full outage.
02Constraints
- One Redis instance, not clustered; 4 vCPUs; Redis is single-threaded for command execution (I/O threads, when enabled, only move bytes across sockets — every command still executes on the one main thread)
- Normal peak: 50,000 reads/sec, spread across ~10M distinct post IDs
- Viral spike: 200,000 reads/sec total, of which 160,000/sec for a single post ID
- Cache TTL is a flat 60s for all post entries — set by the same writer, so every expiry is synchronized with its own writes
- Postgres can sustain ~2,000 reads/sec comfortably before latency degrades
- 30 app servers behind a load balancer; no internal coordination between them (no shared singleflight, no shared backfill lock)
- You cannot predict in advance which post will go viral
- The viral post exists in Redis (it was written normally and served with the rest until its reads exploded)
03Evidence
- Redis command-rate graph during the spike: ~160K of the ~200K ops/sec are GET for the same key (same post ID); the other 40K are spread across the rest of the keyspace
- Redis CPU pinned at 100% the whole window, while memory is fine — no evictions, no OOM, no maxmemory activity
- Redis p99 is uniform across ALL keys during the spike: the slow requests are not the viral key — GETs for cold, unrelated, even nonexistent keys all share the same ~300ms tail
- The instant before the hot key's 60s TTL expires, Postgres CPU is steady (<10%); the second after it expires, Postgres CPU jumps to 100% and every app server logs the same SELECT for the same post ID starting within milliseconds of each other
- APM shows thousands of identical in-flight queries for the same post row, all started in the same second, no request coalescing anywhere in the stack
→The question
Reads for one key at 160K/sec slow down requests for every other key in the same Redis — why, when Redis is otherwise 'fast enough' for the volume? And what exactly breaks in the seconds after that key's TTL hits zero?
04Your prediction
0 / 600 chars