Loss-aware Binarization of Deep Networks
arXiv:1611.01600
Abstract
Deep neural network models, though very powerful and highly successful, are computationally expensive in terms of space and time. Recently, there have been a number of attempts on binarizing the network weights and activations. This greatly reduces the network size, and replaces the underlying multiplications to additions or even XNOR bit operations. However, existing binarization schemes are based on simple matrix approximation and ignore the effect of binarization on the loss. In this paper, we propose a proximal Newton algorithm with diagonal Hessian approximation that directly minimizes the loss w.r.t. the binarized weights. The underlying proximal step has an efficient closed-form solution, and the second-order information can be efficiently obtained from the second moments already computed by the Adam optimizer. Experiments on both feedforward and recurrent networks show that the proposed loss-aware binarization algorithm outperforms existing binarization schemes, and is also more robust for wide and deep networks.
Cited by in corpus (34)
- A Survey of Model Compression and Acceleration for Deep Neural Networks
- Binary Neural Networks: A Survey
- A Survey on Methods and Theories of Quantized Neural Networks
- Loss-aware Weight Quantization of Deep Networks
- Applications and Techniques for Fast Machine Learning in Science
- ProxQuant: Quantized Neural Networks via Proximal Operators
- Rotated Binary Neural Network
- Compacting Deep Neural Networks for Internet of Things: Methods and Applications
- BinaryBERT: Pushing the Limit of BERT Quantization
- Learning Sparse Low-Precision Neural Networks With Learnable Regularization
- Forward and Backward Information Retention for Accurate Binary Neural Networks
- Bolt: Accelerated Data Mining with Fast Vector Compression
- Regularizing Activation Distribution for Training Binarized Deep Networks
- Adaptive Low-Precision Training for Embeddings in Click-Through Rate Prediction
- Structured Binary Neural Networks for Accurate Image Classification and Semantic Segmentation
- BlockSwap: Fisher-guided Block Substitution for Network Compression on a Budget
- FxP-QNet: A Post-Training Quantizer for the Design of Mixed Low-Precision DNNs with Dynamic Fixed-Point Representation
- TernaryBERT: Distillation-aware Ultra-low Bit BERT
- BiPointNet: Binary Neural Network for Point Clouds
- Mirror Descent View for Neural Network Quantization
- Efficient Neural Architecture Search via Proximal Iterations
- Effective Training of Convolutional Neural Networks with Low-bitwidth Weights and Activations
- Bi-Real Net: Binarizing Deep Network Towards Real-Network Performance
- RBCN: Rectified Binary Convolutional Networks for Enhancing the Performance of 1-bit DCNNs
- SinReQ: Generalized Sinusoidal Regularization for Low-Bitwidth Deep Quantized Training
- No Multiplication? No Floating Point? No Problem! Training Networks for Efficient Inference
- Learning Recurrent Binary/Ternary Weights
- Distribution-sensitive Information Retention for Accurate Binary Neural Network
- FALCON: Lightweight and Accurate Convolution
- SiMaN: Sign-to-Magnitude Network Binarization
- Training Quantized Neural Networks with a Full-precision Auxiliary Module
- FeTa: A DCA Pruning Algorithm with Generalization Error Guarantees
- Composite Binary Decomposition Networks
- Sub-bit Neural Networks: Learning to Compress and Accelerate Binary Neural Networks