Why size concurrency against p99 instead of average latency?
Because averages hide the tail. Under load, the slow requests are the ones that pile up and exhaust your connection pool or thread pool. Sizing capacity against p99 (or higher) latency means the system stays healthy when a fraction of requests are slow, which is exactly when it matters.