3 papers
cs.LG2026
ScaleBITS: Scalable Bitwidth Search for Hardware-Aligned Mixed-Precision LLMs
Xinlin Li, Timothy Chou, Josh Fromm +3
Post-training weight quantization is crucial for reducing the memory and inference cost of large language models (LLMs), yet pushing the average precision below 4 bits remains chal…
cs.LG2026
SNIP: An Adaptive Mixed Precision Framework for Subbyte Large Language Model Training
Yunjie Pan, Yongyi Yang, Hanmei Yang +1
Training large language models (LLMs) efficiently while preserving model quality poses significant challenges, particularly with subbyte precision supported by state-of-the-art GPU…
cs.AR2026
Unveiling the Potential of Quantization with MXFP4: Strategies for Quantization Error Reduction
Jatin Chhugani, Geonhwa Jeong, Bor-Yiing Su +8
Large Language Models (LLMs) have intensified the need for low-precision formats that enable efficient, large-scale inference. The Open Compute Project (OCP) Microscaling (MX) stan…