100,000 requests in. The counter says 97,832 — every run, a different shortfall.
01Symptom
A single `counter++` in a Go HTTP handler backs a site-wide hit counter. Under load it is never exact: 100,000 requests produce 97,832 the first run, 98,041 the second. The shortfall changes every run and is never above the true count. No database, no unharnessed resource — just one shared `int` and the default net/http server.
02Constraints
- Go's default `net/http` server: one goroutine per incoming request — nothing pools or serializes handlers
- Load test sends exactly 100,000 requests with high concurrency — thousands in flight at once
- Single machine, multi-core
- No locks, no atomics, no channels: the handler is exactly the shown `counter++`
- The counter must be exactly correct, and the fix is not acceptable without naming the mechanism
03Evidence
- Two identical load runs land on different shortfalls (97,832 then 98,041), and neither ever exceeds 100,000
- `go build -race` reports an unsynchronized read/write on `counter` on the first request batch
- amd64 disassembly of the handler shows three instructions for the increment: a load into a register, an add of one, and a store back — not a single in-memory increment
- Verified at low concurrency (one request at a time) the counter is exact for all 100,000 — the loss scales with overlap, not with volume
→
Recent
Previously diagnosed.
- 80% of reads hit one post ID. Redis pins at 100% and every key's p99 goes 1ms → 300ms. Then the TTL expires.redis · caching · hot-key · thundering-herd · postgreshard
- 98% cache hit rate → 40%, every 60 seconds, like clockworkredis · caching · postgresmedium
- p50 is fine. p99 is 6 seconds. Every endpoint, not just the slow one.postgres · connection-pool · multi-tenantmedium
- CPU graph says 45%. p99 latency says otherwise.jvm · garbage-collection · latencyhard