16 citations · 21 across the 24 of their papers we have counts for
1 paper · 1 filter
Xinrui Chen, Hongxing Zhang, Fanyi Zeng +5
Layer pruning has emerged as a promising technique for compressing large language models (LLMs) while achieving acceleration proportional to the pruning ratio. In this work, we ide…