NICE: Noise Injection and Clamping Estimation for Neural Network Quantization
arXiv:1810.00162 · doi:10.3390/math9172144
Abstract
Convolutional Neural Networks (CNN) are very popular in many fields including computer vision, speech recognition, natural language processing, to name a few. Though deep learning leads to groundbreaking performance in these domains, the networks used are very demanding computationally and are far from real-time even on a GPU, which is not power efficient and therefore does not suit low power systems such as mobile devices. To overcome this challenge, some solutions have been proposed for quantizing the weights and activations of these networks, which accelerate the runtime significantly. Yet, this acceleration comes at the cost of a larger error. The \uniqname method proposed in this work trains quantized neural networks by noise injection and a learned clamping, which improve the accuracy. This leads to state-of-the-art results on various regression and classification tasks, e.g., ImageNet classification with architectures such as ResNet-18/34/50 with low as 3-bit weights and activations. We implement the proposed solution on an FPGA to demonstrate its applicability for low power real-time applications. The implementation of the paper is available at https://github.com/Lancer555/NICE
References in corpus (9)
- Trained Ternary Quantization
- DeepISP: Towards Learning an End-to-End Image Processing Pipeline
- WRPN: Wide Reduced-Precision Networks
- Apprentice: Using Knowledge Distillation Techniques To Improve Low-Precision Network Accuracy
- UNIQ: Uniform Noise Injection for Non-Uniform Quantization of Neural Networks
- Stronger generalization bounds for deep nets via a compression approach
- LQ-Nets: Learned Quantization for Highly Accurate and Compact Deep Neural Networks
- Learning Accurate Low-Bit Deep Neural Networks with Stochastic Quantization
- HCM: Hardware-Aware Complexity Metric for Neural Network Architectures
Cited by in corpus (16)
- Learned Step Size Quantization
- Trained Quantization Thresholds for Accurate and Efficient Fixed-Point Inference of Deep Neural Networks
- Robust Quantization: One Model to Rule Them All
- Pareto-Optimal Quantized ResNet Is Mostly 4-bit
- Differentiable Model Compression via Pseudo Quantization Noise
- FQ-Conv: Fully Quantized Convolution for Efficient and Accurate Inference
- CAT: Compression-Aware Training for bandwidth reduction
- Convolutional Neural Networks Quantization with Attention
- Differentiable Dynamic Quantization with Mixed Precision and Adaptive Resolution
- Towards Learning of Filter-Level Heterogeneous Compression of Convolutional Neural Networks
- Neural gradients are near-lognormal: improved quantized and sparse training
- FAT: Learning Low-Bitwidth Parametric Representation via Frequency-Aware Transformation
- A High-Performance Adaptive Quantization Approach for Edge CNN Applications
- RTN: Reparameterized Ternary Network
- A Very Compact Embedded CNN Processor Design Based on Logarithmic Computing
- Exploring Neural Networks Quantization via Layer-Wise Quantization Analysis