2 papers
cs.AR2026
Unveiling the Potential of Quantization with MXFP4: Strategies for Quantization Error Reduction
Jatin Chhugani, Geonhwa Jeong, Bor-Yiing Su +8
Large Language Models (LLMs) have intensified the need for low-precision formats that enable efficient, large-scale inference. The Open Compute Project (OCP) Microscaling (MX) stan…
cs.AR2025
COMET: A Framework for Modeling Compound Operation Dataflows with Explicit Collectives
Shubham Negi, Manik Singhal, Aayush Ankit +2
Modern machine learning accelerators are designed to efficiently execute deep neural networks (DNNs) by optimizing data movement, memory hierarchy, and compute throughput. However,…