1 paper · 1 filter
Bowen Pang, Kai Li, Ruifeng She +1
With the development of large language models (LLMs), it has become increasingly important to optimize hardware usage and improve throughput. In this paper, we study the inference…