ShiftAddNet: A Hardware-Inspired Deep Network
arXiv:2010.12785
Abstract
Multiplication (e.g., convolution) is arguably a cornerstone of modern deep neural networks (DNNs). However, intensive multiplications cause expensive resource costs that challenge DNNs' deployment on resource-constrained edge devices, driving several attempts for multiplication-less deep networks. This paper presented ShiftAddNet, whose main inspiration is drawn from a common practice in energy-efficient hardware implementation, that is, multiplication can be instead performed with additions and logical bit-shifts. We leverage this idea to explicitly parameterize deep networks in this way, yielding a new type of deep network that involves only bit-shift and additive weight layers. This hardware-inspired ShiftAddNet immediately leads to both energy-efficient inference and training, without compromising the expressive capacity compared to standard DNNs. The two complementary operation types (bit-shift and add) additionally enable finer-grained control of the model's learning capacity, leading to more flexible trade-off between accuracy and (training) efficiency, as well as improved robustness to quantization and pruning. We conduct extensive experiments and ablation studies, all backed up by our FPGA-based ShiftAddNet implementation and energy measurements. Compared to existing DNNs or other multiplication-less models, ShiftAddNet aggressively reduces over 80% hardware-quantified energy cost of DNNs training and inference, while offering comparable or better accuracies. Codes and pre-trained models are available at https://github.com/RICE-EIC/ShiftAddNet.
Accepted by NeurIPS 2020
References in corpus (14)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
- Binarized Neural Networks: Training Deep Neural Networks with Weights and Activations Constrained to +1 or -1
- Rethinking the Value of Network Pruning
- Scalable Methods for 8-bit Training of Neural Networks
- Neural Networks with Few Multiplications
- Latent Weights Do Not Exist: Rethinking Binarized Neural Network Optimization
- ImageNet Training in Minutes
- Kernel Based Progressive Distillation for Adder Neural Networks
- On the Expressive Power of Overlapping Architectures of Deep Learning
- Shift-based Primitives for Efficient Convolutional Neural Networks
- SmartExchange: Trading Higher-cost Memory Storage/Access for Lower-cost Computation
- AdderSR: Towards Energy Efficient Image Super-Resolution
- Orthogonal Over-Parameterized Training
Cited by in corpus (11)
- Artificial Neural Networks for Photonic Applications: From Algorithms to Implementation
- Reducing Computational Complexity of Neural Networks in Optical Channel Equalization: From Concepts to Implementation
- AdderSR: Towards Energy Efficient Image Super-Resolution
- Manifold Regularized Dynamic Network Pruning
- Attention Mechanism with Energy-Friendly Operations
- DietCNN: Multiplication-free Inference for Quantized CNNs
- NASA: Neural Architecture Search and Acceleration for Hardware Inspired Hybrid Networks
- Token Shift Transformer for Video Classification
- "BNN - BN = ?": Training Binary Neural Networks without Batch Normalization
- SmartDeal: Re-Modeling Deep Network Weights for Efficient Inference and Training
- : Sign-Sparse-Shift Reparametrization for Effective Training of Low-bit Shift Networks