Databasespostgresconnection-poolmulti-tenantmedium

p50 is fine. p99 is 6 seconds. Every endpoint, not just the slow one.

01Symptom

A multi-tenant API shares one Postgres connection pool (size 20) across all endpoints. Most queries run in under 10ms. But under load, p99 latency across EVERY endpoint — including trivial ones — balloons to seconds, while Postgres CPU stays low.

02Constraints

  • Single shared pool, 20 connections, no per-endpoint isolation
  • One reporting endpoint runs a genuinely slow query (5-8s) for certain filter combinations, used by a handful of tenants
  • No statement_timeout set on the pool
  • Fast endpoints (simple key lookups) normally return in <10ms

03Evidence

  • Pool 'connections in use' metric pins at 20/20 for sustained periods during the slow windows
  • Fast endpoints' own query execution time (measured inside the handler, after acquiring a connection) is still <10ms — the delay is entirely in pool acquisition wait time
  • The slow report endpoint's traffic volume correlates exactly with the onset of the pool exhaustion windows
  • Postgres server-side CPU and active query count stay low — it's not struggling to execute queries, connections just aren't available to hand out

→The question

Why does a slow endpoint used by a few tenants degrade latency for every tenant on every endpoint?

04Your prediction

01Where's the actual bottleneck?
02What would confirm this before changing any code?
0 / 600 chars