2 citations · 3 across the 4 of their papers we have counts for
1 paper · 2 filters
Songtao Liu, Peng Liu
Pruning is a highly effective approach for compressing large language models (LLMs), significantly reducing inference latency. However, conventional training-free structured prunin…