3 papers
cs.DC2026
Piper: Efficient Large-Scale MoE Training via Resource Modeling and Pipelined Hybrid Parallelism
Sajal Dash, Feiyi Wang
Frontier models increasingly adopt Mixture-of-Experts (MoE) architectures to achieve large-model performance at reduced cost. However, training MoE models on HPC platforms is hinde…
cs.LG2026
A Parallel Alternative for Energy-Efficient Neural Network Training and Inferencing
Sudip K. Seal, Maksudul Alam, Jorge Ramirez +2
Energy efficiency of training and inferencing with large neural network models is a critical challenge facing the future of sustainable large-scale machine learning workloads. This…
cs.LG2025
X-MoE: Enabling Scalable Training for Emerging Mixture-of-Experts Architectures on HPC Platforms
Yueming Yuan, Ahan Gupta, Jianping Li +3
Emerging expert-specialized Mixture-of-Experts (MoE) architectures, such as DeepSeek-MoE, deliver strong model quality through fine-grained expert segmentation and large top-k rout…