2 papers
cs.LG2026
GlowQ: Group-Shared LOw-Rank Approximation for Quantized LLMs
Selim An, Il hong Suh, Yeseong Kim
Quantization techniques such as BitsAndBytes, AWQ, and GPTQ are widely used as a standard method in deploying large language models but often degrades accuracy when using low-bit r…
cs.CL2026
FineScope : SAE-guided Data Selection Enables Domain Specific LLM Pruning and Finetuning
Chaitali Bhattacharyya, Hyunsei Lee, Junyoung Lee +3
Training large language models (LLMs) from scratch requires significant computational resources, driving interest in developing smaller, domain-specific LLMs that maintain both eff…