1 paper
Zejia Lin, Hongxin Xu, Guanyi Chen +3
Modern LLM serving systems confront inefficient GPU utilization due to the fundamental mismatch between compute-intensive prefill and memory-bound decode phases. While current prac…