TensorDash: Exploiting Sparsity to Accelerate Deep Neural Network Training and Inference
arXiv:2009.00748 · doi:10.1109/MICRO50266.2020.00069
Abstract
TensorDash is a hardware level technique for enabling data-parallel MAC units to take advantage of sparsity in their input operand streams. When used to compose a hardware accelerator for deep learning, TensorDash can speedup the training process while also increasing energy efficiency. TensorDash combines a low-cost, sparse input operand interconnect comprising an 8-input multiplexer per multiplier input, with an area-efficient hardware scheduler. While the interconnect allows a very limited set of movements per operand, the scheduler can effectively extract sparsity when it is present in the activations, weights or gradients of neural networks. Over a wide set of models covering various applications, TensorDash accelerates the training process by while being more energy-efficient, more energy efficient when taking on-chip and off-chip memory accesses into account. While TensorDash works with any datatype, we demonstrate it with both single-precision floating-point units and bfloat16.
References in corpus (7)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Practical Bayesian Optimization of Machine Learning Algorithms
- Show and Tell: Lessons learned from the 2015 MSCOCO Image Captioning Challenge
- Training Deep Neural Networks with 8-bit Floating Point Numbers
- Parameter Efficient Training of Deep Convolutional Neural Networks by Dynamic Sparse Reparameterization
- TensorDash: Exploiting Sparsity to Accelerate Deep Neural Network Training and Inference
- A Study of BFLOAT16 for Deep Learning Training
Cited by in corpus (8)
- SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning
- TensorDash: Exploiting Sparsity to Accelerate Deep Neural Network Training and Inference
- TeAAL: A Declarative Framework for Modeling Sparse Tensor Accelerators
- Shift-BNN: Highly-Efficient Probabilistic Bayesian Neural Network Training via Memory-Friendly Pattern Retrieving
- Demystifying BERT: Implications for Accelerator Design
- HASCO: Towards Agile HArdware and Software CO-design for Tensor Computation
- Laplacian Pyramid-like Autoencoder
- HASI: Hardware-Accelerated Stochastic Inference, A Defense Against Adversarial Machine Learning Attacks