p50 is fine. p99 is 6 seconds. Every endpoint, not just the slow one.
01Symptom
A multi-tenant API shares one Postgres connection pool (size 20) across all endpoints. Most queries run in under 10ms. But under load, p99 latency across EVERY endpoint — including trivial ones — balloons to seconds, while Postgres CPU stays low.
02Constraints
- Single shared pool, 20 connections, no per-endpoint isolation
- One reporting endpoint runs a genuinely slow query (5-8s) for certain filter combinations, used by a handful of tenants
- No statement_timeout set on the pool
- Fast endpoints (simple key lookups) normally return in <10ms
03Evidence
- Pool 'connections in use' metric pins at 20/20 for sustained periods during the slow windows
- Fast endpoints' own query execution time (measured inside the handler, after acquiring a connection) is still <10ms — the delay is entirely in pool acquisition wait time
- The slow report endpoint's traffic volume correlates exactly with the onset of the pool exhaustion windows
- Postgres server-side CPU and active query count stay low — it's not struggling to execute queries, connections just aren't available to hand out
→The question
Why does a slow endpoint used by a few tenants degrade latency for every tenant on every endpoint?
04Your prediction
0 / 600 chars