1 paper
Gangmuk Lim, Wanyu Zhao, Brighten Godfrey +3
Efficiently serving large language model (LLM) inference tasks is crucial both for user-perceived latency such as time-to-first-token (TTFT) and for GPU utilization. However, LLM r…