1 citations · 1 across the 1 of their papers we have counts for
1 paper
Zheng Wang, Boxiao Jin, Zhongzhi Yu +1
How to efficiently serve Large Language Models (LLMs) has become a pressing issue because of their huge computational cost in their autoregressive generation process. To mitigate c…