1 citations · 1 across the 1 of their papers we have counts for
1 paper
Yi Zhang, Fei Yang, Shuang Peng +2
Large language models (LLMs) have demonstrated state-of-the-art performance across various tasks. However, the latency of inference and the large GPU memory consumption of LLMs res…