1 paper
Zhixiang Wei, Yun Wang, James Yen +2
LLM inference comprises a compute-bound prefill phase and a memory-bound decode phase, and recent systems disaggregate them onto separate hardware. Yet today's datacenter GPUs rely…