39 citations · 39 across the 3 of their papers we have counts for
3 papers
cs.LG2023★ 39 cited
A Hybrid Tensor-Expert-Data Parallelism Approach to Optimize Mixture-of-Experts Training
Siddharth Singh, Olatunji Ruwase, Ammar Ahmad Awan +3
Mixture-of-Experts (MoE) is a neural network architecture that adds sparsely activated expert blocks to a base model, increasing the number of parameters without impacting computat…
cs.LG2023
Exploiting Sparsity in Pruned Neural Networks to Optimize Large Model Training
Siddharth Singh, Abhinav Bhatele
Parallel training of neural networks at scale is challenging due to significant overheads arising from communication. Recently, deep learning researchers have developed a variety o…
math.GR2022
Certain properties of the enhanced power graph associated with a finite group
Parveen, Jitender Kumar, Siddharth Singh +1
The enhanced power graph of a finite group , denoted by , is the simple undirected graph whose vertex set is and two distinct vertices are adjacent…