1 paper
Junyoung Lee, Sehyeon Park, Shinhyoung Jang +5
Pruning is a practical approach to compress large language models (LLMs), but it can amplify text degeneration, especially repetition loops, even when perplexity and task accuracy…