10 citations · 13 across the 2 of their papers we have counts for
2 papers
cs.DC2025★ 10 cited
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism
Zheng Zhang, Donglin Yang, Yaqi Xia +4
Recently, Mixture-of-Experts (MoE) has become one of the most popular techniques to scale pre-trained models to extraordinarily large sizes. Dynamic activation of experts allows fo…
cs.CL2022★ 3 cited
Vega-MT: The JD Explore Academy Translation System for WMT22
Changtong Zan, Keqin Peng, Liang Ding +9
We describe the JD Explore Academy's submission of the WMT 2022 shared general translation task. We participated in all high-resource tracks and one medium-resource track, includin…