2 papers
cs.AI2026
RaBiT: Residual-Aware Binarization Training for Accurate and Efficient LLMs
Youngcheon You, Banseok Lee, Minseop Choi +5
Efficient deployment of large language models (LLMs) requires extreme quantization, forcing a critical trade-off between low-bit efficiency and performance. Residual binarization e…
cs.LG2026
LittleBit: Ultra Low-Bit Quantization via Latent Factorization
Banseok Lee, Dongkyu Kim, Youngcheon You +1
The deployment of large language models (LLMs) is frequently hindered by prohibitive memory and computational requirements. While quantization mitigates these bottlenecks, maintain…