1 paper · 1 filter
Peiqi Yu, Jinhao Wang, Xinyi Sui +3
Post-training pruning is an effective approach for reducing the size and inference cost of large language models (LLMs), but existing methods often face a trade-off between pruning…