activity
20162020
most citedOn Scale-out Deep Learning Training for Cloud and HPC

19 citations · 44 across the 4 of their papers we have counts for

collaborators

6 papers

cs.NI20203 cited

Machine Learning (ML) In a 5G Standalone (SA) Self Organizing Network (SON)

Srinivasan Sridharan

Machine learning (ML) is included in Self-organizing Networks (SONs) that are key drivers for enhancing the Operations, Administration, and Maintenance (OAM) activities. It is incl…

cs.DC2020

Deep Learning Training in Facebook Data Centers: Design of Scale-up and Scale-out Systems

Maxim Naumov, John Kim, Dheevatsa Mudigere +12

Large-scale training is important to ensure high performance and accuracy of machine-learning models. At Facebook we use many different models, including computer vision, video and…

cs.DC20193 cited

Automatic Model Parallelism for Deep Neural Networks with Compiler and Hardware Support

Sanket Tavarageri, Srinivas Sridharan, Bharat Kaul

The deep neural networks (DNNs) have been enormously successful in tasks that were hitherto in the human-only realm such as image recognition, and language translation. Owing to th…

cs.DC201819 cited

On Scale-out Deep Learning Training for Cloud and HPC

Srinivas Sridharan, Karthikeyan Vaidyanathan, Dhiraj Kalamkar +8

The exponential growth in use of large deep neural networks has accelerated the need for training these deep neural networks in hours or even minutes. This can only be achieved thr…

cs.PF201719 cited

Deep Learning at 15PF: Supervised and Semi-Supervised Classification for Scientific Data

Thorsten Kurth, Jian Zhang, Nadathur Satish +12

This paper presents the first, 15-PetaFLOP Deep Learning system for solving scientific pattern classification problems on contemporary HPC architectures. We develop supervised conv…

cs.DC2016

Distributed Deep Learning Using Synchronous Stochastic Gradient Descent

Dipankar Das, Sasikanth Avancha, Dheevatsa Mudigere +5

We design and implement a distributed multinode synchronous SGD algorithm, without altering hyper parameters, or compressing data, or altering algorithmic behavior. We perform a de…