1 paper · 1 filter
Zhengyan Zhang, Yixin Song, Guanghui Yu +7
Sparse computation offers a compelling solution for the inference of Large Language Models (LLMs) in low-resource scenarios by dynamically skipping the computation of inactive neur…