614 citations · 691 across the 8 of their papers we have counts for
1 paper · 1 filter
Xiteng Yao, Taeho Kim, Hengzhi Pei +7
As large language models (LLMs) move into production serving, practitioners must rapidly evaluate inference performance across diverse hardware, models, and serving parameters to m…