61 citations · 134 across the 6 of their papers we have counts for
7 papers
On Scale-out Deep Learning Training for Cloud and HPC
Srinivas Sridharan, Karthikeyan Vaidyanathan, Dhiraj Kalamkar +8
The exponential growth in use of large deep neural networks has accelerated the need for training these deep neural networks in hours or even minutes. This can only be achieved thr…
Galactos: Computing the Anisotropic 3-Point Correlation Function for 2 Billion Galaxies
Brian Friesen, Md. Mostofa Ali Patwary, Brian Austin +8
The nature of dark energy and the complete theory of gravity are two central questions currently facing cosmology. A vital tool for addressing them is the 3-point correlation funct…
Deep Learning at 15PF: Supervised and Semi-Supervised Classification for Scientific Data
Thorsten Kurth, Jian Zhang, Nadathur Satish +12
This paper presents the first, 15-PetaFLOP Deep Learning system for solving scientific pattern classification problems on contemporary HPC architectures. We develop supervised conv…
Ternary Residual Networks
Abhisek Kundu, Kunal Banerjee, Naveen Mellempudi +4
Sub-8-bit representation of DNNs incur some discernible loss of accuracy despite rigorous (re)training at low-precision. Such loss of accuracy essentially makes them equivalent to…
Ternary Neural Networks with Fine-Grained Quantization
Naveen Mellempudi, Abhisek Kundu, Dheevatsa Mudigere +3
We propose a novel fine-grained quantization (FGQ) method to ternarize pre-trained full precision models, while also constraining activations to 8 and 4-bits. Using this method, we…
Distributed Deep Learning Using Synchronous Stochastic Gradient Descent
Dipankar Das, Sasikanth Avancha, Dheevatsa Mudigere +5
We design and implement a distributed multinode synchronous SGD algorithm, without altering hyper parameters, or compressing data, or altering algorithmic behavior. We perform a de…