2 citations · 5 across the 9 of their papers we have counts for
1 paper · 1 filter
Zijian Feng, Hanzhang Zhou, Zixiao Zhu +5
Pruning is a widely used technique to reduce the size and inference cost of large language models (LLMs), but it often causes performance degradation. To mitigate this, existing re…