3 papers
cs.LG2026
LoKA: Low-precision Kernel Applications for Recommendation Models At Scale
Liang Luo, Yinbin Ma, Quanyu Zhu +21
Recent GPU generations deliver significantly higher FLOPs using lower-precision arithmetic, such as FP8. While successfully applied to large language models (LLMs), its adoption in…
cs.PF2026
Optimus: A Generic Operator-Level PyTorch Model Transformation Framework
Menglu Yu, Jiaqi Xu, Yuzhen Huang +19
In large-scale industrial applications, deep learning models that power recommendation and ranking have complex and diverse model architectures. These models are continuously devel…
cs.DC2025
Two-dimensional Sparse Parallelism for Large Scale Deep Learning Recommendation Model Training
Xin Zhang, Quanyu Zhu, Liangbei Xu +8
The increasing complexity of deep learning recommendation models (DLRM) has led to a growing need for large-scale distributed systems that can efficiently train vast amounts of dat…