1 paper
Ziyan Wang, Enmao Diao, Qi Le +6
Structured pruning is a practical approach to deploying large language models (LLMs) efficiently, as it yields compact, hardware-friendly architectures. However, the dominant local…