39 citations · 39 across the 3 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2023★ 39 cited
A Hybrid Tensor-Expert-Data Parallelism Approach to Optimize Mixture-of-Experts Training
Siddharth Singh, Olatunji Ruwase, Ammar Ahmad Awan +3
Mixture-of-Experts (MoE) is a neural network architecture that adds sparsely activated expert blocks to a base model, increasing the number of parameters without impacting computat…
cs.LG2023
Exploiting Sparsity in Pruned Neural Networks to Optimize Large Model Training
Siddharth Singh, Abhinav Bhatele
Parallel training of neural networks at scale is challenging due to significant overheads arising from communication. Recently, deep learning researchers have developed a variety o…