1 paper · 1 filter
Mengting He, Shihao Xia, Haomin Jia +2
The widespread adoption of large language models (LLMs) has made GPU-accelerated inference a critical part of modern computing infrastructure. Production inference systems rely on…