5 papers
Unleashing Low-Bit Inference on Ascend NPUs: A Comprehensive Evaluation of HiFloat Formats
Pengxiang Zhao, Hui-Ling Zhen, Xing Li +10
As LLMs scale, low-bit floating-point formats like MXFP and NVFP4 offer new opportunities for precision and efficiency. In this work, we evaluate HiFloat (HiF8 and HiF4), a family…
ButterflyQuant: Ultra-low-bit LLM Quantization through Learnable Orthogonal Butterfly Transforms
Bingxin Xu, Zhen Dong, Oussama Elachqar +1
Large language models require massive memory footprints, severely limiting deployment on consumer hardware. Quantization reduces memory through lower numerical precision, but extre…
Transformer-Based Tooth Alignment Prediction With Occlusion And Collision Constraints
ZhenXing Dong, JiaZhou Chen, YangHui Xu
The planning of digital orthodontic treatment requires providing tooth alignment, which not only consumes a lot of time and labor to determine manually but also relays clinical exp…
Stochastic Communication Avoidance for Recommendation Systems
Lutfi Eren Erdogan, Vijay Anand Raghava Kanakagiri, Kurt Keutzer +1
One of the major bottlenecks for efficient deployment of neural network based recommendation systems is the memory footprint of their embedding tables. Although many neural network…
DQRM: Deep Quantized Recommendation Models
Yang Zhou, Zhen Dong, Ellick Chan +3
Large-scale recommendation models are currently the dominant workload for many large Internet companies. These recommenders are characterized by massive embedding tables that are s…