1 paper
Sicheng Xu, Hao Shi, Wei Zhang +6
Structured pruning is a promising approach for reducing the substantial inference costs of Large Language Models (LLMs) while maintaining hardware efficiency. Many existing methods…