Fixed Point Quantization of Deep Convolutional Networks
arXiv:1511.06393
Abstract
In recent years increasingly complex architectures for deep convolution networks (DCNs) have been proposed to boost the performance on image recognition tasks. However, the gains in performance have come at a cost of substantial increase in computation and model storage resources. Fixed point implementation of DCNs has the potential to alleviate some of these complexities and facilitate potential deployment on embedded hardware. In this paper, we propose a quantizer design for fixed point implementation of DCNs. We formulate and solve an optimization problem to identify optimal fixed point bit-width allocation across DCN layers. Our experiments show that in comparison to equal bit-width settings, the fixed point DCNs with optimized bit width allocation offer >20% reduction in the model size without any loss in accuracy on CIFAR-10 benchmark. We also demonstrate that fine-tuning can further enhance the accuracy of fixed point DCNs beyond that of the original floating point model. In doing so, we report a new state-of-the-art fixed point performance of 6.78% error-rate on CIFAR-10 benchmark.
ICML 2016
References in corpus (10)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Deep Residual Learning for Image Recognition
- Going Deeper with Convolutions
- Deep Learning with Limited Numerical Precision
- Compressing Deep Convolutional Networks using Vector Quantization
- Return of the Devil in the Details: Delving Deep into Convolutional Nets
- Compressing Neural Networks with the Hashing Trick
- XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks
- Neural Networks with Few Multiplications
- Overcoming Challenges in Fixed Point Training of Deep Convolutional Networks
Cited by in corpus (142)
- Enabling Spike-based Backpropagation for Training Deep Neural Network Architectures
- Point-Voxel CNN for Efficient 3D Deep Learning
- NullHop: A Flexible Convolutional Neural Network Accelerator Based on Sparse Representations of Feature Maps
- CirCNN: Accelerating and Compressing Deep Neural Networks Using Block-CirculantWeight Matrices
- An Energy-Efficient FPGA-based Deconvolutional Neural Networks Accelerator for Single Image Super-Resolution
- Hardware Acceleration of Sparse and Irregular Tensor Computations of ML Models: A Survey and Insights
- Rethinking floating point for deep learning
- Layer-specific Optimization for Mixed Data Flow with Mixed Precision in FPGA Design for CNN-based Object Detectors
- Taurus: A Data Plane Architecture for Per-Packet ML
- Deep Neural Network Approximation for Custom Hardware: Where We've Been, Where We're Going
- Improving Post Training Neural Quantization: Layer-wise Calibration and Integer Programming
- Making AI Forget You: Data Deletion in Machine Learning
- AddNet: Deep Neural Networks Using FPGA-Optimized Multipliers
- Rotated Binary Neural Network
- Deep Learning with Low Precision by Half-wave Gaussian Quantization
- Compact recurrent neural networks for acoustic event detection on low-energy low-complexity platforms
- Up or Down? Adaptive Rounding for Post-Training Quantization
- QKD: Quantization-aware Knowledge Distillation
- Improved training of binary networks for human pose estimation and image recognition
- Hierarchical binary CNNs for landmark localization with limited resources
- Improving Efficiency in Convolutional Neural Network with Multilinear Filters
- Convolutional-Recurrent Neural Networks on Low-Power Wearable Platforms for Cardiac Arrhythmia Detection
- Searching for Low-Bit Weights in Quantized Neural Networks
- ExPAN(N)D: Exploring Posits for Efficient Artificial Neural Network Design in FPGA-based Systems
- BinaryDuo: Reducing Gradient Mismatch in Binary Activation Network by Coupling Binary Activations
- Overcoming Challenges in Fixed Point Training of Deep Convolutional Networks
- Non-Structured DNN Weight Pruning -- Is It Beneficial in Any Platform?
- Memory-Driven Mixed Low Precision Quantization For Enabling Deep Network Inference On Microcontrollers
- Feature Map Transform Coding for Energy-Efficient CNN Inference
- Progressive DNN Compression: A Key to Achieve Ultra-High Weight Pruning and Quantization Rates using ADMM
- Acceleration of Convolutional Neural Network Using FFT-Based Split Convolutions
- Learned Threshold Pruning
- Quantune: Post-training Quantization of Convolutional Neural Networks using Extreme Gradient Boosting for Fast Deployment
- Resource-Efficient Neural Networks for Embedded Systems
- AutoQ: Automated Kernel-Wise Neural Network Quantization
- Q-CapsNets: A Specialized Framework for Quantizing Capsule Networks
- Regularizing Activation Distribution for Training Binarized Deep Networks
- A transprecision floating-point cluster for efficient near-sensor data analytics
- Fixed-point Quantization of Convolutional Neural Networks for Quantized Inference on Embedded Platforms
- Efficient Execution of Quantized Deep Learning Models: A Compiler Approach
- Check-N-Run: A Checkpointing System for Training Deep Learning Recommendation Models
- Deep Learning as a Mixed Convex-Combinatorial Optimization Problem
- To compress or not to compress: Understanding the Interactions between Adversarial Attacks and Neural Network Compression
- Quantization of Deep Neural Networks for Accumulator-constrained Processors
- Precision Gating: Improving Neural Network Efficiency with Dynamic Dual-Precision Activations
- Post-Training Piecewise Linear Quantization for Deep Neural Networks
- Binary Ensemble Neural Network: More Bits per Network or More Networks per Bit?
- The ZipML Framework for Training Models with End-to-End Low Precision: The Cans, the Cannots, and a Little Bit of Deep Learning
- REQ-YOLO: A Resource-Aware, Efficient Quantization Framework for Object Detection on FPGAs
- Toward Extremely Low Bit and Lossless Accuracy in DNNs with Progressive ADMM
- EffNet: An Efficient Structure for Convolutional Neural Networks
- CoCoPIE: Making Mobile AI Sweet As PIE --Compression-Compilation Co-Design Goes a Long Way
- DaCapo: Accelerating Continuous Learning in Autonomous Systems for Video Analytics
- Trained Rank Pruning for Efficient Deep Neural Networks
- Low-Precision Reinforcement Learning: Running Soft Actor-Critic in Half Precision
- Rethinking Differentiable Search for Mixed-Precision Neural Networks
- ADMM-NN: An Algorithm-Hardware Co-Design Framework of DNNs Using Alternating Direction Method of Multipliers
- Automated Design Space Exploration for optimised Deployment of DNN on Arm Cortex-A CPUs
- In-Ear-Voice: Towards Milli-Watt Audio Enhancement With Bone-Conduction Microphones for In-Ear Sensing Platforms
- Per-Tensor Fixed-Point Quantization of the Back-Propagation Algorithm
- CacheNet: A Model Caching Framework for Deep Learning Inference on the Edge
- A Heterogeneous Parallel Non-von Neumann Architecture System for Accurate and Efficient Machine Learning Molecular Dynamics
- Iterative Low-Rank Approximation for CNN Compression
- Boosting Binary Masks for Multi-Domain Learning through Affine Transformations
- Bit Error Robustness for Energy-Efficient DNN Accelerators
- An Application-Specific VLIW Processor with Vector Instruction Set for CNN Acceleration
- Lightweight Convolutional Representations for On-Device Natural Language Processing
- SYMOG: learning symmetric mixture of Gaussian modes for improved fixed-point quantization
- Toward Compact Deep Neural Networks via Energy-Aware Pruning
- S2DNAS:Transforming Static CNN Model for Dynamic Inference via Neural Architecture Search
- An Automotive Case Study on the Limits of Approximation for Object Detection
- Term Revealing: Furthering Quantization at Run Time on Quantized DNNs
- Adaptive Precision Training: Quantify Back Propagation in Neural Networks with Fixed-point Numbers
- Scalar Quantization as Sparse Least Square Optimization
- DARC: Differentiable ARchitecture Compression
- A flexible FPGA accelerator for convolutional neural networks
- Phoenix: A Low-Precision Floating-Point Quantization Oriented Architecture for Convolutional Neural Networks
- End-to-End Learned Image Compression with Quantized Weights and Activations
- A Simple Method to Reduce Off-chip Memory Accesses on Convolutional Neural Networks
- A Unified DNN Weight Compression Framework Using Reweighted Optimization Methods
- A High-Performance Adaptive Quantization Approach for Edge CNN Applications
- Training for 'Unstable' CNN Accelerator:A Case Study on FPGA
- LCP: A Low-Communication Parallelization Method for Fast Neural Network Inference in Image Recognition
- Matrix and tensor decompositions for training binary neural networks
- Approximations in Deep Learning
- Fixed-Point Code Synthesis For Neural Networks
- Neural Machine Translation with 4-Bit Precision and Beyond
- Confounding Tradeoffs for Neural Network Quantization
- Optimal Quantization for Batch Normalization in Neural Network Deployments and Beyond
- Softmax Tempering for Training Neural Machine Translation Models
- Towards Explainable Bit Error Tolerance of Resistive RAM-Based Binarized Neural Networks
- MAFAT: Memory-Aware Fusing and Tiling of Neural Networks for Accelerated Edge Inference
- Table-Based Neural Units: Fully Quantizing Networks for Multiply-Free Inference
- MUSCO: Multi-Stage Compression of neural networks
- SGQuant: Squeezing the Last Bit on Graph Neural Networks with Specialized Quantization
- A Very Compact Embedded CNN Processor Design Based on Logarithmic Computing
- ResOT: Resource-Efficient Oblique Trees for Neural Signal Classification
- Q-Rater: Non-Convex Optimization for Post-Training Uniform Quantization
- Learning In Practice: Reasoning About Quantization
- Encrypted Speech Recognition using Deep Polynomial Networks
- Knowledge Distillation For Recurrent Neural Network Language Modeling With Trust Regularization
- E-RNN: Design Optimization for Efficient Recurrent Neural Networks in FPGAs
- Balancing Cost and Benefit with Tied-Multi Transformers
- WaLDORf: Wasteless Language-model Distillation On Reading-comprehension
- NGEMM: Optimizing GEMM for Deep Learning via Compiler-based Techniques
- EDEN: Enabling Energy-Efficient, High-Performance Deep Neural Network Inference Using Approximate DRAM
- Helix: Algorithm/Architecture Co-design for Accelerating Nanopore Genome Base-calling
- Reducing Inference Latency with Concurrent Architectures for Image Recognition
- Irregularly Tabulated MLP for Fast Point Feature Embedding
- Verifying Quantized Neural Networks using SMT-Based Model Checking
- Training Quantized Neural Networks to Global Optimality via Semidefinite Programming
- FATNN: Fast and Accurate Ternary Neural Networks
- Separating the Effects of Batch Normalization on CNN Training Speed and Stability Using Classical Adaptive Filter Theory
- An Inter-Layer Weight Prediction and Quantization for Deep Neural Networks based on a Smoothly Varying Weight Hypothesis
- Efficient and Robust Machine Learning for Real-World Systems
- Towards Accurate and High-Speed Spiking Neuromorphic Systems with Data Quantization-Aware Deep Networks
- Barrier-Free Large-Scale Sparse Tensor Accelerator (BARISTA) For Convolutional Neural Networks
- Smoothed Differential Privacy
- BitSplit-Net: Multi-bit Deep Neural Network with Bitwise Activation Function
- KML: Using Machine Learning to Improve Storage Systems
- A Highly Effective Low-Rank Compression of Deep Neural Networks with Modified Beam-Search and Modified Stable Rank
- Joint Coreset Construction and Quantization for Distributed Machine Learning
- Recurrent Stacking of Layers in Neural Networks: An Application to Neural Machine Translation
- Quantized Neural Networks via {-1, +1} Encoding Decomposition and Acceleration
- GRIM: A General, Real-Time Deep Learning Inference Framework for Mobile Devices based on Fine-Grained Structured Weight Sparsity
- Closed-Loop Neural Prostheses with On-Chip Intelligence: A Review and A Low-Latency Machine Learning Model for Brain State Detection
- Seeing Convolution Through the Eyes of Finite Transformation Semigroup Theory: An Abstract Algebraic Interpretation of Convolutional Neural Networks
- Understanding and Overcoming the Challenges of Efficient Transformer Quantization
- Phantom: A High-Performance Computational Core for Sparse Convolutional Neural Networks
- Booster: An Accelerator for Gradient Boosting Decision Trees
- An Overview of Datatype Quantization Techniques for Convolutional Neural Networks
- XpulpNN: Enabling Energy Efficient and Flexible Inference of Quantized Neural Network on RISC-V based IoT End Nodes
- Design and Analysis of Uplink and Downlink Communications for Federated Learning
- Multi-Precision Quantized Neural Networks via Encoding Decomposition of -1 and +1
- Multi-Task Network Pruning and Embedded Optimization for Real-time Deployment in ADAS
- Training Multi-bit Quantized and Binarized Networks with A Learnable Symmetric Quantizer
- PERMDNN: Efficient Compressed DNN Architecture with Permuted Diagonal Matrices
- Towards Modality Transferable Visual Information Representation with Optimal Model Compression
- IFQ-Net: Integrated Fixed-point Quantization Networks for Embedded Vision
- Closed-Loop Neural Interfaces with Embedded Machine Learning
- When deep learning models on GPU can be accelerated by taking advantage of unstructured sparsity
- Robust error bounds for quantised and pruned neural networks