1 paper · 1 filter
Weiqiao Shan, Long Meng, Tong Zheng +5
Large language models (LLMs) exhibit exceptional performance across various downstream tasks. However, they encounter limitations due to slow inference speeds stemming from their e…