1 paper
Jae Gon Kim, Donghoon Yoo, Hanyul Ryu +3
Datacenter GPU power is the binding constraint on LLM serving capacity, and production serving has shifted to prefill/decode (PD) disaggregation. Deploying NVIDIA's Max-Q inference…