1 paper
Jaewon Cheon, Pilsung Kang
The growing size of large language models has created significant computational inefficiencies. To address this challenge, sparse activation methods selectively deactivates non-ess…