1 paper
Junyi Wu, Chao Fang, Zhongfeng Wang
The growing scale of large language models (LLMs) has intensified demands on computation and memory, making efficient inference a key challenge. While sparsity can reduce these cos…