32 citations · 35 across the 2 of their papers we have counts for
2 papers
cs.DC2022★ 32 cited
Tutel: Adaptive Mixture-of-Experts at Scale
Changho Hwang, Wei Cui, Yifan Xiong +12
Sparsely-gated mixture-of-experts (MoE) has been widely adopted to scale deep learning models to trillion-plus parameters with fixed computational cost. The algorithmic performance…
cs.DC2022★ 3 cited
GC3: An Optimizing Compiler for GPU Collective Communication
Meghan Cowan, Saeed Maleki, Madanlal Musuvathi +2
Machine learning models made up of millions or billions of parameters are trained and served on large multi-GPU systems. As models grow in size and execute on more GPUs, the collec…