3 papers
cs.DC2026
Direct Model State Migration for Elastic Training of Large Language Models
Weijian Liu, Mingzhen Li, Rui Kang +3
Large language model (LLM) training shall adapt to dynamic resources in shared clusters to tackle the elasticity, including passive preemption and optimistic scaling. State migrati…
cs.DC2026
JanusPipe: Efficient Pipeline Parallel Training for Machine Learning Interatomic Potentials
Hongyu Wang, Weijian Liu, Hongtao Xu +4
Discovering atom-level phenomena requires molecular dynamics (MD) simulations with ab initio accuracy. Machine learning interatomic potentials (MLIPs) enable stable, high-accuracy…
cs.DC2026
Breaking the Training Barrier of Billion-Parameter Universal Machine Learning Interatomic Potentials
Yuanchang Zhou, Hongyu Wang, Yiming Du +12
Universal Machine Learning Interatomic Potentials (uMLIPs), pre-trained on massively diverse datasets encompassing inorganic materials and organic molecules across the entire perio…