2 citations · 2 across the 5 of their papers we have counts for
Showing cs.DCShow all
3 papers · 1 filter
cs.DC2026
LAER-MoE: Load-Adaptive Expert Re-layout for Efficient Mixture-of-Experts Training
Xinyi Liu, Yujie Wang, Fangcheng Fu +4
Expert parallelism is vital for effectively training Mixture-of-Experts (MoE) models, enabling different devices to host distinct experts, with each device processing different inp…
cs.DC2025
Galvatron: An Automatic Distributed System for Efficient Foundation Model Training
Xinyi Liu, Yujie Wang, Shenhan Zhu +4
Galvatron is a distributed system for efficiently training large-scale Foundation Models. It overcomes the complexities of selecting optimal parallelism strategies by automatically…
cs.DC2024★ 2 cited
FlexSP: Accelerating Large Language Model Training via Flexible Sequence Parallelism
Yujie Wang, Shiju Wang, Shenhan Zhu +7
Extending the context length (i.e., the maximum supported sequence length) of LLMs is of paramount significance. To facilitate long context training of LLMs, sequence parallelism h…