Runtimegoconnection-poolconcurrencypprofmedium

5,000 req/sec. App server CPU at 100%, Postgres under 20%. Latency goes from 5ms to 400ms.

01Symptom

A URL shortener: one Go process, one Postgres, GET /r/{shortcode} looks up a long URL and issues a redirect. Under a light load it is fine at 5ms. Under a load test at 5,000 requests/sec, the app server's CPU climbs to 100% while Postgres CPU stays under 20% — and per-request latency balloons to 400ms.

02Constraints

  • Single Go process, 4 CPU cores, 8GB RAM
  • Single Postgres instance on a separate machine, mostly idle (<20% CPU) during the incident
  • shortcode has a B-tree index; the query is exactly SELECT long_url FROM urls WHERE shortcode = $1
  • Each DB query takes ~2ms round-trip including network
  • Sustained 5,000 req/sec, one DB query per request
  • Go's database/sql pool settings are untouched — no MaxOpenConns, no MaxIdleConns, no lifetime
  • No caching layer in front of the database

03Evidence

  • Postgres CPU never rises, and the query is a single indexed lookup on an equality match — the database is not the constraint
  • The handler is trivial: one query, one redirect, almost no work of its own
  • The process is running far more concurrent requests than 4 cores can execute, all of them blocked on the same small set of DB connections
  • No pprof endpoint was exposed, so nobody has a profile of where the CPU is actually going

→The question

If Postgres is idle, what is burning 100% of the app server's CPU, and what would you check before changing a single line of code?

04Your prediction

01Where is the CPU actually going?
02What do you do first?
03What is the actual fix?
0 / 600 chars