9 citations · 9 across the 1 of their papers we have counts for
2 papers
cs.DC2024★ 1 cited
Lancet: Accelerating Mixture-of-Experts Training via Whole Graph Computation-Communication Overlapping
Chenyu Jiang, Ye Tian, Zhen Jia +3
The Mixture-of-Expert (MoE) technique plays a crucial role in expanding the size of DNN model parameters. However, it faces the challenge of extended all-to-all communication laten…
cs.DC2023★ 9 cited
DynaPipe: Optimizing Multi-task Training through Dynamic Pipelines
Chenyu Jiang, Zhen Jia, Shuai Zheng +2
Multi-task model training has been adopted to enable a single deep neural network model (often a large language model) to handle multiple tasks (e.g., question answering and text s…