8 papers
Libra: Taming Attention Workload Skew in Long-Context LLM Training with Bounded Sequence Pool
Yan Wang, Xiulong Yuan, Kaiming Yang +16
Long-context LLM training suffers from a load-balancing problem that sequence packing does not solve. Packing samples into fixed-token sequences balances memory and linear-cost ope…
JanusPipe: Efficient Pipeline Parallel Training for Machine Learning Interatomic Potentials
Hongyu Wang, Weijian Liu, Hongtao Xu +4
Discovering atom-level phenomena requires molecular dynamics (MD) simulations with ab initio accuracy. Machine learning interatomic potentials (MLIPs) enable stable, high-accuracy…
SparseBalance: Load-Balanced Long Context Training with Dynamic Sparse Attention
Hongtao Xu, Jianchao Tan, Yuxuan Hu +8
While sparse attention mitigates the computational bottleneck of long-context LLM training, its distributed training process exhibits extreme heterogeneity in both \textit{1)} sequ…
Breaking the Training Barrier of Billion-Parameter Universal Machine Learning Interatomic Potentials
Yuanchang Zhou, Hongyu Wang, Yiming Du +12
Universal Machine Learning Interatomic Potentials (uMLIPs), pre-trained on massively diverse datasets encompassing inorganic materials and organic molecules across the entire perio…
MatRIS: Toward Reliable and Efficient Pretrained Machine Learning Interatomic Potentials
Yuanchang Zhou, Siyu Hu, Xiangyu Zhang +3
Foundation MLIPs demonstrate broad applicability across diverse material systems and have emerged as a powerful and transformative paradigm in chemical and computational materials…
fix pimd/langevin: An Efficient Implementation of Path Integral Molecular Dynamics in LAMMPS
Yifan Li, Axel Gomez, Kehan Cai +10
Path integral molecular dynamics (PIMD), which maps a quantum particle onto a fictitious classical system of ring polymers and propagates the "beads" of this extended classical sys…