7 citations · 15 across the 3 of their papers we have counts for
3 papers
Efficient distributed algorithms for Convolutional Neural Networks
Rui Li, Yufan Xu, Aravind Sukumaran-Rajam +2
Several efficient distributed algorithms have been developed for matrix-matrix multiplication: the 3D algorithm, the 2D SUMMA algorithm, and the 2.5D algorithm. Each of these algor…
PL-NMF: Parallel Locality-Optimized Non-negative Matrix Factorization
Gordon E. Moon, Aravind Sukumaran-Rajam, Srinivasan Parthasarathy +1
Non-negative Matrix Factorization (NMF) is a key kernel for unsupervised dimension reduction used in a wide range of applications, including topic modeling, recommender systems and…
Load-Balanced Sparse MTTKRP on GPUs
Israt Nisa, Jiajia Li, Aravind Sukumaran-Rajam +2
Sparse matricized tensor times Khatri-Rao product (MTTKRP) is one of the most computationally expensive kernels in sparse tensor computations. This work focuses on optimizing the M…