1 paper
Yidu Wu, Xiang Wang, Kejie Zhao +3
Large language models (LLMs) achieve strong generation and reasoning performance, but the Transformer architecture incurs high inference cost. Existing acceleration methods often r…