Neural Networks with Few Multiplications
arXiv:1510.03009
Abstract
For most deep learning algorithms training is notoriously time consuming. Since most of the computation in training neural networks is typically spent on floating point multiplications, we investigate an approach to training that eliminates the need for most of these. Our method consists of two parts: First we stochastically binarize weights to convert multiplications involved in computing hidden states to sign changes. Second, while back-propagating error derivatives, in addition to binarizing the weights, we quantize the representations at each layer to convert the remaining multiplications into binary shifts. Experimental results across 3 popular datasets (MNIST, CIFAR10, SVHN) show that this approach not only does not hurt classification performance but can result in even better performance than standard stochastic gradient descent training, paving the way to fast, hardware-friendly training of neural networks.
Published as a conference paper at ICLR 2016. 9 pages, 3 figures
References in corpus (5)
- BinaryConnect: Training Deep Neural Networks with binary weights during propagations
- On Using Monolingual Corpora in Neural Machine Translation
- Adding Gradient Noise Improves Learning for Very Deep Networks
- Bitwise Neural Networks
- Training Binary Multilayer Neural Networks for Image Classification using Expectation Backpropagation
Cited by in corpus (53)
- Binarized Neural Networks: Training Deep Neural Networks with Weights and Activations Constrained to +1 or -1
- DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients
- A Survey of Model Compression and Acceleration for Deep Neural Networks
- Ternary Weight Networks
- Trained Ternary Quantization
- Fixed Point Quantization of Deep Convolutional Networks
- Convolutional Neural Networks using Logarithmic Data Representation
- Accelerating Eulerian Fluid Simulation With Convolutional Networks
- XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks
- A Survey on Methods and Theories of Quantized Neural Networks
- WRPN: Wide Reduced-Precision Networks
- Flexpoint: An Adaptive Numerical Format for Efficient Training of Deep Neural Networks
- Edge Intelligence: Architectures, Challenges, and Applications
- Mixed-precision deep learning based on computational memory
- Loss-aware Weight Quantization of Deep Networks
- Reduced-Precision Strategies for Bounded Memory in Deep Neural Nets
- Deep Neural Network Approximation for Custom Hardware: Where We've Been, Where We're Going
- Applications and Techniques for Fast Machine Learning in Science
- Deep Learning with Low Precision by Half-wave Gaussian Quantization
- Towards the Limit of Network Quantization
- Structured Probabilistic Pruning for Convolutional Neural Network Acceleration
- On-Device Machine Learning: An Algorithms and Learning Theory Perspective
- ShiftAddNet: A Hardware-Inspired Deep Network
- Compressing RNNs for IoT devices by 15-38x using Kronecker Products
- Overcoming Challenges in Fixed Point Training of Deep Convolutional Networks
- Run-Time Efficient RNN Compression for Inference on Edge Devices
- Packing Sparse Convolutional Neural Networks for Efficient Systolic Array Implementations: Column Combining Under Joint Optimization
- Binary Ensemble Neural Network: More Bits per Network or More Networks per Bit?
- REQ-YOLO: A Resource-Aware, Efficient Quantization Framework for Object Detection on FPGAs
- LCNN: Lookup-based Convolutional Neural Network
- ASP Vision: Optically Computing the First Layer of Convolutional Neural Networks using Angle Sensitive Pixels
- Structured Pruning for Efficient ConvNets via Incremental Regularization
- Boosting Binary Masks for Multi-Domain Learning through Affine Transformations
- Understanding the Impact of Precision Quantization on the Accuracy and Energy of Neural Networks
- GXNOR-Net: Training deep neural networks with ternary weights and activations without full-precision memory under a unified discretization framework
- BitNet: Bit-Regularized Deep Neural Networks
- Phoenix: A Low-Precision Floating-Point Quantization Oriented Architecture for Convolutional Neural Networks
- Binary Complex Neural Network Acceleration on FPGA
- Training for 'Unstable' CNN Accelerator:A Case Study on FPGA
- Generalized Ternary Connect: End-to-End Learning and Compression of Multiplication-Free Deep Neural Networks
- Detecting Dead Weights and Units in Neural Networks
- Deep counter networks for asynchronous event-based processing
- Optimal Quantization for Batch Normalization in Neural Network Deployments and Beyond
- Recent Advances in Efficient Computation of Deep Convolutional Neural Networks
- Smaller Models, Better Generalization
- Efficient and Robust Machine Learning for Real-World Systems
- Pruning Ternary Quantization
- Fixed-point Factorized Networks
- DecisiveNets: Training Deep Associative Memories to Solve Complex Machine Learning Problems
- EEG Channel Interpolation Using Deep Encoder-decoder Netwoks
- : Sign-Sparse-Shift Reparametrization for Effective Training of Low-bit Shift Networks
- Adaptive Binary-Ternary Quantization
- Low Precision Floating-point Arithmetic for High Performance FPGA-based CNN Acceleration