September 2026 · vLLM · A100 · GPU Profiling
Beyond Peak Throughput: Finding and Explaining vLLM’s Concurrency Knee on an A100
A controlled saturation study showing how increasing concurrency beyond the throughput knee can reduce throughput, sharply worsen latency, and move GPU execution into a more expensive mixed/prefill attention regime.
Read investigation →