21 citations · 22 across the 4 of their papers we have counts for
1 paper · 1 filter
Yi Zhang, Fei Yang, Shuang Peng +2
Large language models (LLMs) have demonstrated state-of-the-art performance across various tasks. However, the latency of inference and the large GPU memory consumption of LLMs res…