2 papers
cs.LG2026
NanoQuant: Efficient Sub-1-Bit Quantization of Large Language Models
Hyochan Chong, Dongkyu Kim, Changdong Kim +1
Weight-only quantization has become a standard approach for efficiently serving large language models (LLMs). However, existing methods fail to efficiently compress models to binar…
cs.AI2026
RaBiT: Residual-Aware Binarization Training for Accurate and Efficient LLMs
Youngcheon You, Banseok Lee, Minseop Choi +5
Efficient deployment of large language models (LLMs) requires extreme quantization, forcing a critical trade-off between low-bit efficiency and performance. Residual binarization e…