66 citations · 299 across the 20 of their papers we have counts for
4 papers · 1 filter
Extending Sparse Tensor Accelerators to Support Multiple Compression Formats
Eric Qin, Geonhwa Jeong, William Won +7
Sparsity, which occurs in both scientific applications and Deep Learning (DL) models, has been a key target of optimization within recent ASIC accelerators due to the potential mem…
High Performance Scalable FPGA Accelerator for Deep Neural Networks
Sudarshan Srinivasan, Pradeep Janedula, Saurabh Dhoble +7
Low-precision is the first order knob for achieving higher Artificial Intelligence Operations (AI-TOPS). However the algorithmic space for sub-8-bit precision compute is diverse, w…
On Scale-out Deep Learning Training for Cloud and HPC
Srinivas Sridharan, Karthikeyan Vaidyanathan, Dhiraj Kalamkar +8
The exponential growth in use of large deep neural networks has accelerated the need for training these deep neural networks in hours or even minutes. This can only be achieved thr…
Distributed Deep Learning Using Synchronous Stochastic Gradient Descent
Dipankar Das, Sasikanth Avancha, Dheevatsa Mudigere +5
We design and implement a distributed multinode synchronous SGD algorithm, without altering hyper parameters, or compressing data, or altering algorithmic behavior. We perform a de…