1 paper · 1 filter
Tianteng Gu, Bei Liu, Bo Xiao +3
Pruning is a widely used technique to compress large language models (LLMs) by removing unimportant weights, but it often suffers from significant performance degradation - especia…