39 citations · 39 across the 2 of their papers we have counts for
3 papers
cs.LG2022
A Roadmap for Big Model
Sha Yuan, Hanyu Zhao, Shuai Zhao +97
With the rapid development of deep learning, training Big Models (BMs) for multiple downstream tasks becomes a popular paradigm. Researchers have achieved various outcomes in the c…
cs.LG2021★ 39 cited
FastMoE: A Fast Mixture-of-Expert Training System
Jiaao He, Jiezhong Qiu, Aohan Zeng +3
Mixture-of-Expert (MoE) presents a strong potential in enlarging the size of language model to trillions of parameters. However, training trillion-scale MoE requires algorithm and…
cs.DC2019
Heterogeneity-Aware Asynchronous Decentralized Training
Qinyi Luo, Jiaao He, Youwei Zhuo +1
Distributed deep learning training usually adopts All-Reduce as the synchronization mechanism for data parallel algorithms due to its high performance in homogeneous environment. H…