01
Archive
10 incidents, newest first.
10 of 10
- 100,000 requests in. The counter says 97,832 — every run, a different shortfall.go · concurrency · race-condition · atomics · runtimemedium
- 80% of reads hit one post ID. Redis pins at 100% and every key's p99 goes 1ms → 300ms. Then the TTL expires.redis · caching · hot-key · thundering-herd · postgreshard
- 98% cache hit rate → 40%, every 60 seconds, like clockworkredis · caching · postgresmedium
- p50 is fine. p99 is 6 seconds. Every endpoint, not just the slow one.postgres · connection-pool · multi-tenantmedium
- CPU graph says 45%. p99 latency says otherwise.jvm · garbage-collection · latencyhard
- Traffic spikes. Pods start dying. Healthy pods, killed by their own cluster.kubernetes · cascading-failure · load-sheddinghard
- No deploy. No traffic change. Queries just keep getting slower, day over day.postgres · autovacuum · bloatmedium
- Support is full of tickets. Some customers were charged twice, and some were charged once but have no order.idempotency · payments · retries · distributed-systemshard
- 40 rows in 1ms. 6 million rows in 9 seconds, with a sequential scan — on a table that has an index on user_id.query-planner · indexes · sequential-scan · mvcchard
- 5,000 req/sec. App server CPU at 100%, Postgres under 20%. Latency goes from 5ms to 400ms.go · connection-pool · concurrency · pprofmedium