Quantizing deep convolutional networks for efficient inference: A whitepaper
arXiv:1806.08342
Abstract
We present an overview of techniques for quantizing convolutional neural networks for inference with integer weights and activations. Per-channel quantization of weights and per-layer quantization of activations to 8-bits of precision post-training produces classification accuracies within 2% of floating point networks for a wide variety of CNN architectures. Model sizes can be reduced by a factor of 4 by quantizing weights to 8-bits, even when 8-bit arithmetic is not supported. This can be achieved with simple, post training quantization of weights.We benchmark latencies of quantized networks on CPUs and DSPs and observe a speedup of 2x-3x for quantized implementations compared to floating point on CPUs. Speedups of up to 10x are observed on specialized processors with fixed point SIMD capabilities, like the Qualcomm QDSPs with HVX. Quantization-aware training can provide further improvements, reducing the gap to floating point to 1% at 8-bit precision. Quantization-aware training also allows for reducing the precision of weights to four bits with accuracy losses ranging from 2% to 10%, with higher accuracy drop for smaller networks.We introduce tools in TensorFlow and TensorFlowLite for quantizing convolutional networks and review best practices for quantization-aware training to obtain high accuracy with quantized weights and activations. We recommend that per-channel quantization of weights and per-layer quantization of activations be the preferred quantization scheme for hardware acceleration and kernel optimization. We also propose that future processors and hardware accelerators for optimized inference support precisions of 4, 8 and 16 bits.
37 pages
Cited by in corpus (155)
- Efficient Deep Learning: A Survey on Making Deep Learning Models Smaller, Faster, and Better
- Binary Neural Networks: A Survey
- Machine Learning for Microcontroller-Class Hardware: A Review
- Integer Quantization for Deep Learning Inference: Principles and Empirical Evaluation
- TensorFlow Lite Micro: Embedded Machine Learning on TinyML Systems
- MicroNets: Neural Network Architectures for Deploying TinyML Applications on Commodity Microcontrollers
- Hardware Approximate Techniques for Deep Neural Network Accelerators: A Survey
- Post-training 4-bit quantization of convolution networks for rapid-deployment
- Training with Quantization Noise for Extreme Model Compression
- Hardware Acceleration of Sparse and Irregular Tensor Computations of ML Models: A Survey and Insights
- Visual Wake Words Dataset
- Deep Learning Inference in Facebook Data Centers: Characterization, Performance Optimizations and Hardware Implications
- Bayesian Bits: Unifying Quantization and Pruning
- Improving Post Training Neural Quantization: Layer-wise Calibration and Integer Programming
- An Electro-Photonic System for Accelerating Deep Neural Networks
- Differentiable Soft Quantization: Bridging Full-Precision and Low-Bit Neural Networks
- PLAM: a Posit Logarithm-Approximate Multiplier
- Trained Quantization Thresholds for Accurate and Efficient Fixed-Point Inference of Deep Neural Networks
- Q-SpiNN: A Framework for Quantizing Spiking Neural Networks
- HAWQ-V2: Hessian Aware trace-Weighted Quantization of Neural Networks
- AdaPT: Fast Emulation of Approximate DNN Accelerators in PyTorch
- Q-BERT: Hessian Based Ultra Low Precision Quantization of BERT
- Neural Network Distiller: A Python Package For DNN Compression Research
- Up or Down? Adaptive Rounding for Post-Training Quantization
- Pruning and Quantization for Deep Neural Network Acceleration: A Survey
- HAQ: Hardware-Aware Automated Quantization with Mixed Precision
- QKD: Quantization-aware Knowledge Distillation
- RAELLA: Reforming the Arithmetic for Efficient, Low-Resolution, and Low-Loss Analog PIM: No Retraining Required!
- K for the Price of 1: Parameter-efficient Multi-task and Transfer Learning
- ReLeQ: A Reinforcement Learning Approach for Deep Quantization of Neural Networks
- EasyQuant: Post-training Quantization via Scale Optimization
- HAWQ: Hessian AWare Quantization of Neural Networks with Mixed-Precision
- Robust Quantization: One Model to Rule Them All
- Towards Efficient Training for Neural Network Quantization
- Lightweight Compression of Intermediate Neural Network Features for Collaborative Intelligence
- DSConv: Efficient Convolution Operator
- Efficient Winograd Convolution via Integer Arithmetic
- Memory-Driven Mixed Low Precision Quantization For Enabling Deep Network Inference On Microcontrollers
- Rethinking "Batch" in BatchNorm
- COIN: COmpression with Implicit Neural representations
- EnforceSNN: Enabling Resilient and Energy-Efficient Spiking Neural Network Inference considering Approximate DRAMs for Embedded Systems
- Quantune: Post-training Quantization of Convolutional Neural Networks using Extreme Gradient Boosting for Fast Deployment
- Computer Vision Model Compression Techniques for Embedded Systems: A Survey
- Fighting Quantization Bias With Bias
- Degree-Quant: Quantization-Aware Training for Graph Neural Networks
- AutoQ: Automated Kernel-Wise Neural Network Quantization
- Towards Unified INT8 Training for Convolutional Neural Network
- Fixed-point Quantization of Convolutional Neural Networks for Quantized Inference on Embedded Platforms
- ZeroQ: A Novel Zero Shot Quantization Framework
- Differentiable Model Compression via Pseudo Quantization Noise
- Efficient Execution of Quantized Deep Learning Models: A Compiler Approach
- HAWQV3: Dyadic Neural Network Quantization
- LSQ+: Improving low-bit quantization through learnable offsets and better initialization
- Hardware-Centric AutoML for Mixed-Precision Quantization
- Post-Training Piecewise Linear Quantization for Deep Neural Networks
- Reliability-Aware Quantization for Anti-Aging NPUs
- I-BERT: Integer-only BERT Quantization
- Post-Training 4-bit Quantization on Embedding Tables
- Taxonomy and Evaluation of Structured Compression of Convolutional Neural Networks
- Searching for Winograd-aware Quantized Networks
- A Data and Compute Efficient Design for Limited-Resources Deep Learning
- AI on the Edge: Rethinking AI-based IoT Applications Using Specialized Edge Architectures
- Layer Pruning via Fusible Residual Convolutional Block for Deep Neural Networks
- BatchQuant: Quantized-for-all Architecture Search with Robust Quantizer
- KDLSQ-BERT: A Quantized Bert Combining Knowledge Distillation with Learned Step Size Quantization
- VS-Quant: Per-vector Scaled Quantization for Accurate Low-Precision Neural Network Inference
- A White Paper on Neural Network Quantization
- lpSpikeCon: Enabling Low-Precision Spiking Neural Network Processing for Efficient Unsupervised Continual Learning on Autonomous Agents
- Bit Efficient Quantization for Deep Neural Networks
- Automated Design Space Exploration for optimised Deployment of DNN on Arm Cortex-A CPUs
- A Lite Distributed Semantic Communication System for Internet of Things
- FxP-QNet: A Post-Training Quantizer for the Design of Mixed Low-Precision DNNs with Dynamic Fixed-Point Representation
- CNN2Gate: Toward Designing a General Framework for Implementation of Convolutional Neural Networks on FPGA
- Real-time Denoising and Dereverberation with Tiny Recurrent U-Net
- Gradient Regularization for Quantization Robustness
- Channel-wise Hessian Aware trace-Weighted Quantization of Neural Networks
- SQuantizer: Simultaneous Learning for Both Sparse and Low-precision Neural Networks
- A Generalized Zero-Shot Quantization of Deep Convolutional Neural Networks via Learned Weights Statistics
- INT8 Winograd Acceleration for Conv1D Equipped ASR Models Deployed on Mobile Devices
- Bit Error Robustness for Energy-Efficient DNN Accelerators
- SYMOG: learning symmetric mixture of Gaussian modes for improved fixed-point quantization
- Revisiting Block-based Quantisation: What is Important for Sub-8-bit LLM Inference?
- Rethinking Floating Point Overheads for Mixed Precision DNN Accelerators
- Training Deep Neural Networks with Joint Quantization and Pruning of Weights and Activations
- Integer-only Quantized Transformers for Embedded FPGA-based Time-series Forecasting in AIoT
- At-Scale Evaluation of Weight Clustering to Enable Energy-Efficient Object Detection
- Deep Learning Interference Cancellation in Wireless Networks
- Differentiable Dynamic Quantization with Mixed Precision and Adaptive Resolution
- Towards Real-Time Neural Video Codec for Cross-Platform Application Using Calibration Information
- Training Deep Neural Networks Using Posit Number System
- Post-training Quantization with Multiple Points: Mixed Precision without Mixed Precision
- Generative Zero-shot Network Quantization
- Popcorn: Paillier Meets Compression For Efficient Oblivious Neural Network Inference
- Do All MobileNets Quantize Poorly? Gaining Insights into the Effect of Quantization on Depthwise Separable Convolutional Networks Through the Eyes of Multi-scale Distributional Dynamics
- CRAFT: Criticality-Aware Fault-Tolerance Enhancement Techniques for Emerging Memories-Based Deep Neural Networks
- Artificial neural networks condensation: A strategy to facilitate adaption of machine learning in medical settings by reducing computational burden
- FlowPrecision: Advancing FPGA-Based Real-Time Fluid Flow Estimation with Linear Quantization
- Mantis: Enabling Energy-Efficient Autonomous Mobile Agents with Spiking Neural Networks
- Differentiable Joint Pruning and Quantization for Hardware Efficiency
- Learning low-precision neural networks without Straight-Through Estimator(STE)
- Neural gradients are near-lognormal: improved quantized and sparse training
- Hessian-Aware Pruning and Optimal Neural Implant
- A High-Performance Adaptive Quantization Approach for Edge CNN Applications
- ML-EXray: Visibility into ML Deployment on the Edge
- Compiling Neural Networks for a Computational Memory Accelerator
- FAT: Learning Low-Bitwidth Parametric Representation via Frequency-Aware Transformation
- A 71.2-W Speech Recognition Accelerator with Recurrent Spiking Neural Network
- Network Pruning using Adaptive Exemplar Filters
- IntraQ: Learning Synthetic Images with Intra-Class Heterogeneity for Zero-Shot Network Quantization
- Table-Based Neural Units: Fully Quantizing Networks for Multiply-Free Inference
- Confounding Tradeoffs for Neural Network Quantization
- NeRV: Neural Representations for Videos
- A 1.6-mW Sparse Deep Learning Accelerator for Speech Separation
- Fine-grained Data Distribution Alignment for Post-Training Quantization
- On the Effects of Quantisation on Model Uncertainty in Bayesian Neural Networks
- EASTER: Efficient and Scalable Text Recognizer
- Optimal Quantization for Batch Normalization in Neural Network Deployments and Beyond
- FrostNet: Towards Quantization-Aware Network Architecture Search
- An FPGA Accelerated Method for Training Feed-forward Neural Networks Using Alternating Direction Method of Multipliers and LSMR
- Generative Design of Hardware-aware DNNs
- Non-Blocking Simultaneous Multithreading: Embracing the Resiliency of Deep Neural Networks
- Learned Variable-Rate Image Compression with Residual Divisive Normalization
- LiMuSE: Lightweight Multi-modal Speaker Extraction
- Quantization of Acoustic Model Parameters in Automatic Speech Recognition Framework
- Post-Training Sparsity-Aware Quantization
- Weight Equalizing Shift Scaler-Coupled Post-training Quantization
- Post-Training BatchNorm Recalibration
- SoftNeuro: Fast Deep Inference using Multi-platform Optimization
- A Convergence Theory Towards Practical Over-parameterized Deep Neural Networks
- PTQ-SL: Exploring the Sub-layerwise Post-training Quantization
- Semi-Streaming Architecture: A New Design Paradigm for CNN Implementation on FPGAs
- Pre-trained Language Model based Ranking in Baidu Search
- Towards Sustainable Artificial Intelligence: A Comprehensive Review and Comparative Analysis of Deep Learning Models' Carbon Footprint
- NGEMM: Optimizing GEMM for Deep Learning via Compiler-based Techniques
- Is In-Domain Data Really Needed? A Pilot Study on Cross-Domain Calibration for Network Quantization
- CoDeNet: Efficient Deployment of Input-Adaptive Object Detection on Embedded FPGAs
- Information-Theoretic Understanding of Population Risk Improvement with Model Compression
- MixMix: All You Need for Data-Free Compression Are Feature and Data Mixing
- Conditional Neural Architecture Search
- Full-stack Optimization for Accelerating CNNs with FPGA Validation
- Multi-Task Network Pruning and Embedded Optimization for Real-time Deployment in ADAS
- HPTQ: Hardware-Friendly Post Training Quantization
- Accelerating Neural Network Inference by Overflow Aware Quantization
- Compressing Deep Convolutional Neural Networks by Stacking Low-dimensional Binary Convolution Filters
- Compact retail shelf segmentation for mobile deployment
- An Underexplored Dilemma between Confidence and Calibration in Quantized Neural Networks
- Faster Convolution Inference Through Using Pre-Calculated Lookup Tables
- Low-Precision Hardware Architectures Meet Recommendation Model Inference at Scale
- 4-bit Quantization of LSTM-based Speech Recognition Models
- Auto-Split: A General Framework of Collaborative Edge-Cloud AI
- Understanding and Overcoming the Challenges of Efficient Transformer Quantization
- Memory-Efficient CNN Accelerator Based on Interlayer Feature Map Compression
- Adaptive Precision Training for Resource Constrained Devices
- Hybrid and Non-Uniform quantization methods using retro synthesis data for efficient inference
- Learned Multi-Resolution Variable-Rate Image Compression with Octave-based Residual Blocks