Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations
arXiv:1609.07061
Abstract
We introduce a method to train Quantized Neural Networks (QNNs) --- neural networks with extremely low precision (e.g., 1-bit) weights and activations, at run-time. At train-time the quantized weights and activations are used for computing the parameter gradients. During the forward pass, QNNs drastically reduce memory size and accesses, and replace most arithmetic operations with bit-wise operations. As a result, power consumption is expected to be drastically reduced. We trained QNNs over the MNIST, CIFAR-10, SVHN and ImageNet datasets. The resulting QNNs achieve prediction accuracy comparable to their 32-bit counterparts. For example, our quantized version of AlexNet with 1-bit weights and 2-bit activations achieves top-1 accuracy. Moreover, we quantize the parameter gradients to 6-bits as well which enables gradients computation using only bit-wise operation. Quantized recurrent neural networks were tested over the Penn Treebank dataset, and achieved comparable accuracy as their 32-bit counterparts using only 4-bits. Last but not least, we programmed a binary matrix multiplication GPU kernel with which it is possible to run our MNIST QNN 7 times faster than with an unoptimized GPU kernel, without suffering any loss in classification accuracy. The QNN code is available online.
arXiv admin note: text overlap with arXiv:1602.02830
References in corpus (5)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Deep Learning with Limited Numerical Precision
- Compressing Deep Convolutional Networks using Vector Quantization
- One weird trick for parallelizing convolutional neural networks
- Recurrent Neural Networks With Limited Numerical Precision
Cited by in corpus (120)
- SPINN: Synergistic Progressive Inference of Neural Networks over Device and Cloud
- Hardware and Software Optimizations for Accelerating Deep Neural Networks: Survey of Current Trends, Challenges, and the Road Ahead
- Improving Neural Network Quantization without Retraining using Outlier Channel Splitting
- Scaling Deep Learning on GPU and Knights Landing clusters
- Ternary Neural Networks with Fine-Grained Quantization
- Compacting Deep Neural Networks for Internet of Things: Methods and Applications
- Towards Neural Mixture Recommender for Long Range Dependent User Sequences
- Extending the limit of molecular dynamics with ab initio accuracy to 10 billion atoms
- Mixed Precision Training With 8-bit Floating Point
- Knowledge Distillation via Route Constrained Optimization
- Controlling Information Capacity of Binary Neural Network
- Weight-Sharing Neural Architecture Search: A Battle to Shrink the Optimization Gap
- Knowledge Distillation from Internal Representations
- Optimizing Bit-Serial Matrix Multiplication for Reconfigurable Computing
- QGAN: Quantized Generative Adversarial Networks
- Ultra-compact Binary Neural Networks for Human Activity Recognition on RISC-V Processors
- Projection Convolutional Neural Networks for 1-bit CNNs via Discrete Back Propagation
- Quantization of Deep Neural Networks for Accumulator-constrained Processors
- Automatic Mixed-Precision Quantization Search of BERT
- FQ-Conv: Fully Quantized Convolution for Efficient and Accurate Inference
- Additive Noise Annealing and Approximation Properties of Quantized Neural Networks
- Optimal checkpointing for heterogeneous chains: how to train deep neural networks with limited memory
- DeFINE: DEep Factorized INput Token Embeddings for Neural Sequence Modeling
- EMPIR: Ensembles of Mixed Precision Deep Networks for Increased Robustness against Adversarial Attacks
- Benchmarking Quantized Neural Networks on FPGAs with FINN
- PruneNet: Channel Pruning via Global Importance
- All You Need is a Few Shifts: Designing Efficient Convolutional Neural Networks for Image Classification
- Compression of Acoustic Event Detection Models with Low-rank Matrix Factorization and Quantization Training
- Communication-Efficient Decentralized Learning with Sparsification and Adaptive Peer Selection
- FxP-QNet: A Post-Training Quantizer for the Design of Mixed Low-Precision DNNs with Dynamic Fixed-Point Representation
- CGaP: Continuous Growth and Pruning for Efficient Deep Learning
- On-FPGA Training with Ultra Memory Reduction: A Low-Precision Tensor Method
- A Distributed Synchronous SGD Algorithm with Global Top- Sparsification for Low Bandwidth Networks
- GAN Slimming: All-in-One GAN Compression by A Unified Optimization Framework
- Lightweight Convolutional Representations for On-Device Natural Language Processing
- Improving Adversarial Robustness in Weight-quantized Neural Networks
- An Energy-efficient Time-domain Analog VLSI Neural Network Processor Based on a Pulse-width Modulation Approach
- Boosting Binary Masks for Multi-Domain Learning through Affine Transformations
- Rethinking Floating Point Overheads for Mixed Precision DNN Accelerators
- Training Deep Neural Networks with Joint Quantization and Pruning of Weights and Activations
- Compressibility Loss for Neural Network Weights
- Term Revealing: Furthering Quantization at Run Time on Quantized DNNs
- Efficient Hybrid Network Architectures for Extremely Quantized Neural Networks Enabling Intelligence at the Edge
- Weight Pruning via Adaptive Sparsity Loss
- Topology of deep neural networks
- SoK: How Robust is Image Classification Deep Neural Network Watermarking? (Extended Version)
- Pufferfish: Communication-efficient Models At No Extra Cost
- LP-3DCNN: Unveiling Local Phase in 3D Convolutional Neural Networks
- S-SGD: Symmetrical Stochastic Gradient Descent with Weight Noise Injection for Reaching Flat Minima
- Computation on Sparse Neural Networks: an Inspiration for Future Hardware
- BARS: Joint Search of Cell Topology and Layout for Accurate and Efficient Binary ARchitectures
- Mixed-Precision Quantized Neural Network with Progressively Decreasing Bitwidth For Image Classification and Object Detection
- Generative Zero-shot Network Quantization
- Training Energy-Efficient Deep Spiking Neural Networks with Time-to-First-Spike Coding
- Approximations in Deep Learning
- How Large a Vocabulary Does Text Classification Need? A Variational Approach to Vocabulary Selection
- Knowledge distillation for optimization of quantized deep neural networks
- Pre-trained Language Model for Web-scale Retrieval in Baidu Search
- Compact representations of convolutional neural networks via weight pruning and quantization
- Distributed Low Precision Training Without Mixed Precision
- Progressive Learning of Low-Precision Networks
- Correlation Congruence for Knowledge Distillation
- Ultra-Lightweight Speech Separation via Group Communication
- Sampling-Free Learning of Bayesian Quantized Neural Networks
- Optimal Quantization for Batch Normalization in Neural Network Deployments and Beyond
- An FPGA Accelerated Method for Training Feed-forward Neural Networks Using Alternating Direction Method of Multipliers and LSMR
- Deep Learning in Memristive Nanowire Networks
- ESPN: Extremely Sparse Pruned Networks
- Integer-Only Neural Network Quantization Scheme Based on Shift-Batch-Normalization
- The Pitfall of Evaluating Performance on Emerging AI Accelerators
- WaveQ: Gradient-Based Deep Quantization of Neural Networks through Sinusoidal Adaptive Regularization
- A Layer-wise Adversarial-aware Quantization Optimization for Improving Robustness
- Recurrent Convolution for Compact and Cost-Adjustable Neural Networks: An Empirical Study
- Representation Edit Distance as a Measure of Novelty
- Distributed Learning on Heterogeneous Resource-Constrained Devices
- Spintronics for neuromorphic computing
- GradFreeBits: Gradient Free Bit Allocation for Dynamic Low Precision Neural Networks
- Disentangling Neural Architectures and Weights: A Case Study in Supervised Classification
- EDEN: Enabling Energy-Efficient, High-Performance Deep Neural Network Inference Using Approximate DRAM
- Knowledge Representing: Efficient, Sparse Representation of Prior Knowledge for Knowledge Distillation
- Enhancing efficiency of object recognition in different categorization levels by reinforcement learning in modular spiking neural networks
- Universal Approximation Theorems of Fully Connected Binarized Neural Networks
- TOCO: A Framework for Compressing Neural Network Models Based on Tolerance Analysis
- Quantized Neural Network Inference with Precision Batching
- Training DNNs in O(1) memory with MEM-DFA using Random Matrices
- An Efficient Method of Training Small Models for Regression Problems with Knowledge Distillation
- Analytical aspects of non-differentiable neural networks
- Scalable Smartphone Cluster for Deep Learning
- Finding Everything within Random Binary Networks
- A Distributed SGD Algorithm with Global Sketching for Deep Learning Training Acceleration
- Efficient Inference via Universal LSH Kernel
- Enabling Lightweight Fine-tuning for Pre-trained Language Model Compression based on Matrix Product Operators
- Verifying Quantized Neural Networks using SMT-Based Model Checking
- Learning In Practice: Reasoning About Quantization
- SpotPatch: Parameter-Efficient Transfer Learning for Mobile Object Detection
- A Tapered Floating Point Extension for the Redundant Signed Radix 2 System Using the Canonical Recoding
- The Low-Resource Double Bind: An Empirical Study of Pruning for Low-Resource Machine Translation
- VC dimension of partially quantized neural networks in the overparametrized regime
- Magnitude and Uncertainty Pruning Criterion for Neural Networks
- 4-bit Quantization of LSTM-based Speech Recognition Models
- Hardware realisation of nonlinear dynamical systems for and from biology
- Identifying and Exploiting Structures for Reliable Deep Learning
- Low Power In-Memory Implementation of Ternary Neural Networks with Resistive RAM-Based Synapse
- Modulated binary cliquenet
- Differentiable Architecture Pruning for Transfer Learning
- NN2CAM: Automated Neural Network Mapping for Multi-Precision Edge Processing on FPGA-Based Cameras
- BAMSProd: A Step towards Generalizing the Adaptive Optimization Methods to Deep Binary Model
- Knowledge Distillation By Sparse Representation Matching
- Quantized Neural Networks via {-1, +1} Encoding Decomposition and Acceleration
- Context-Free TextSpotter for Real-Time and Mobile End-to-End Text Detection and Recognition
- Multi-Precision Quantized Neural Networks via Encoding Decomposition of -1 and +1
- Enabling Incremental Training with Forward Pass for Edge Devices
- Maximin Optimization for Binary Regression
- XpulpNN: Enabling Energy Efficient and Flexible Inference of Quantized Neural Network on RISC-V based IoT End Nodes
- Compression of Acoustic Event Detection Models With Quantized Distillation
- Graphs for deep learning representations
- Adaptive Precision Training for Resource Constrained Devices
- Direct Quantization for Training Highly Accurate Low Bit-width Deep Neural Networks
- Demystifying and Generalizing BinaryConnect
- Stochastic Computing for Hardware Implementation of Binarized Neural Networks