10 citations · 12 across the 3 of their papers we have counts for
3 papers
cs.LG2023★ 2 cited
TA-MoE: Topology-Aware Large Scale Mixture-of-Expert Training
Chang Chen, Min Li, Zhihua Wu +2
Sparsely gated Mixture-of-Expert (MoE) has demonstrated its effectiveness in scaling up deep neural networks to an extreme scale. Despite that numerous efforts have been made to im…
cs.CL2020★ 10 cited
Progressively Stacking 2.0: A Multi-stage Layerwise Training Method for BERT Training Speedup
Cheng Yang, Shengnan Wang, Chao Yang +3
Pre-trained language models, such as BERT, have achieved significant accuracy gain in many natural language processing tasks. Despite its effectiveness, the huge number of paramete…
cs.CL2020
CoRe: An Efficient Coarse-refined Training Framework for BERT
Cheng Yang, Shengnan Wang, Yuechuan Li +4
In recent years, BERT has made significant breakthroughs on many natural language processing tasks and attracted great attentions. Despite its accuracy gains, the BERT model genera…