2 papers
cs.LG2026
LoKA: Low-precision Kernel Applications for Recommendation Models At Scale
Liang Luo, Yinbin Ma, Quanyu Zhu +21
Recent GPU generations deliver significantly higher FLOPs using lower-precision arithmetic, such as FP8. While successfully applied to large language models (LLMs), its adoption in…
cs.AR2026
Unveiling the Potential of Quantization with MXFP4: Strategies for Quantization Error Reduction
Jatin Chhugani, Geonhwa Jeong, Bor-Yiing Su +8
Large Language Models (LLMs) have intensified the need for low-precision formats that enable efficient, large-scale inference. The Open Compute Project (OCP) Microscaling (MX) stan…