3 papers
cs.CL2026
FOCUS & RePAIR: Mitigating Text Degeneration via Token-Level Guidance for Pruned Large Language Models
Junyoung Lee, Sehyeon Park, Shinhyoung Jang +5
Pruning is a practical approach to compress large language models (LLMs), but it can amplify text degeneration, especially repetition loops, even when perplexity and task accuracy…
cs.LG2026
GlowQ: Group-Shared LOw-Rank Approximation for Quantized LLMs
Selim An, Il hong Suh, Yeseong Kim
Quantization techniques such as BitsAndBytes, AWQ, and GPTQ are widely used as a standard method in deploying large language models but often degrades accuracy when using low-bit r…
cs.CL2025
FineScope : SAE-guided Data Selection Enables Domain Specific LLM Pruning and Finetuning
Chaitali Bhattacharyya, Hyunsei Lee, Junyoung Lee +3
Training large language models (LLMs) from scratch requires significant computational resources, driving interest in developing smaller, domain-specific LLMs that maintain both eff…