1 paper
Peiqi Yu, Nam Ling, Wei Wang +1
Structured pruning reduces the size and inference cost of large language models (LLMs) by removing weight columns, but the resulting output error can degrade accuracy. Existing tra…