activity
20162026
most citedA Study of BFLOAT16 for Deep Learning Training

66 citations · 237 across the 15 of their papers we have counts for

collaborators
Showing cs.DCShow all

5 papers · 1 filter

cs.DC2020

PolyDL: Polyhedral Optimizations for Creation of High Performance DL primitives

Sanket Tavarageri, Alexander Heinecke, Sasikanth Avancha +3

Deep Neural Networks (DNNs) have revolutionized many aspects of our lives. The use of DNNs is becoming ubiquitous including in softwares for image recognition, speech recognition,…

cs.DC2019

High Performance Scalable FPGA Accelerator for Deep Neural Networks

Sudarshan Srinivasan, Pradeep Janedula, Saurabh Dhoble +7

Low-precision is the first order knob for achieving higher Artificial Intelligence Operations (AI-TOPS). However the algorithmic space for sub-8-bit precision compute is diverse, w…

cs.DC20193 cited

Automatic Model Parallelism for Deep Neural Networks with Compiler and Hardware Support

Sanket Tavarageri, Srinivas Sridharan, Bharat Kaul

The deep neural networks (DNNs) have been enormously successful in tasks that were hitherto in the human-only realm such as image recognition, and language translation. Owing to th…

cs.DC201819 cited

On Scale-out Deep Learning Training for Cloud and HPC

Srinivas Sridharan, Karthikeyan Vaidyanathan, Dhiraj Kalamkar +8

The exponential growth in use of large deep neural networks has accelerated the need for training these deep neural networks in hours or even minutes. This can only be achieved thr…

cs.DC2016

Distributed Deep Learning Using Synchronous Stochastic Gradient Descent

Dipankar Das, Sasikanth Avancha, Dheevatsa Mudigere +5

We design and implement a distributed multinode synchronous SGD algorithm, without altering hyper parameters, or compressing data, or altering algorithmic behavior. We perform a de…