2 papers
cs.DC2024
DiffusionPipe: Training Large Diffusion Models with Efficient Pipelines
Ye Tian, Zhen Jia, Ziyue Luo +2
Diffusion models have emerged as dominant performers for image generation. To support training large diffusion models, this paper studies pipeline parallel training of diffusion mo…
cs.DC2024
Lancet: Accelerating Mixture-of-Experts Training via Whole Graph Computation-Communication Overlapping
Chenyu Jiang, Ye Tian, Zhen Jia +3
The Mixture-of-Expert (MoE) technique plays a crucial role in expanding the size of DNN model parameters. However, it faces the challenge of extended all-to-all communication laten…