1 paper
Zhexuan Gu, Zixun Fu, Yancheng Yuan
Feed-forward networks (FFNs) dominate the parameter count and computation of modern language models, yet existing pruning methods often struggle to convert sparsity into hardware-f…