1 paper
Rui Wen, Lu Sun, Jiayang Liu +3
Compressing large language models reduces memory use and inference cost, but it can also create failures that standard benchmarks miss. A pruned model may still perform well on mul…