1 paper · 1 filter
Rui Wen, Lu Sun, Jiayang Liu +3
Compressing large language models reduces memory use and inference cost, but it can also create failures that standard benchmarks miss. A pruned model may still perform well on mul…