Trained Ternary Quantization
arXiv:1612.01064
Abstract
Deep neural networks are widely used in machine learning applications. However, the deployment of large neural networks models can be difficult to deploy on mobile devices with limited power budgets. To solve this problem, we propose Trained Ternary Quantization (TTQ), a method that can reduce the precision of weights in neural networks to ternary values. This method has very little accuracy degradation and can even improve the accuracy of some models (32, 44, 56-layer ResNet) on CIFAR-10 and AlexNet on ImageNet. And our AlexNet model is trained from scratch, which means it's as easy as to train normal full precision model. We highlight our trained quantization method that can learn both ternary values and ternary assignment. During inference, only ternary values (2-bit weights) and scaling factors are needed, therefore our models are nearly 16x smaller than full-precision models. Our ternary models can also be viewed as sparse binary weight networks, which can potentially be accelerated with custom circuit. Experiments on CIFAR-10 show that the ternary models obtained by trained quantization method outperform full-precision models of ResNet-32,44,56 by 0.04%, 0.16%, 0.36%, respectively. On ImageNet, our model outperforms full-precision AlexNet model by 0.3% of Top-1 accuracy and outperforms previous ternary models by 3%.
Accepted for Poster Presentation on ICLR 2017
References in corpus (1)
Cited by in corpus (67)
- Integer Quantization for Deep Learning Inference: Principles and Empirical Evaluation
- WRPN: Wide Reduced-Precision Networks
- MLPerf Training Benchmark
- Apprentice: Using Knowledge Distillation Techniques To Improve Low-Precision Network Accuracy
- Visual Wake Words Dataset
- Deep Neural Network Approximation for Custom Hardware: Where We've Been, Where We're Going
- DyNet: Dynamic Convolution for Accelerating Convolutional Neural Networks
- Ternary Neural Networks with Fine-Grained Quantization
- QKD: Quantization-aware Knowledge Distillation
- Back to Simplicity: How to Train Accurate BNNs from Scratch?
- Defensive Quantization: When Efficiency Meets Robustness
- A Learning Framework for n-bit Quantized Neural Networks toward FPGAs
- Towards Efficient Training for Neural Network Quantization
- Learning Accurate Low-Bit Deep Neural Networks with Stochastic Quantization
- Mixed Low-precision Deep Learning Inference using Dynamic Fixed Point
- Self-Binarizing Networks
- APQ: Joint Search for Network Architecture, Pruning and Quantization Policy
- QGAN: Quantized Generative Adversarial Networks
- STEERAGE: Synthesis of Neural Networks Using Architecture Search and Grow-and-Prune Methods
- ChamNet: Towards Efficient Network Design through Platform-Aware Model Adaptation
- Learned Low Precision Graph Neural Networks
- Hardware-Centric AutoML for Mixed-Precision Quantization
- FQ-Conv: Fully Quantized Convolution for Efficient and Accurate Inference
- Adaptive Precision CNN Accelerator Using Radix-X Parallel Connected Memristor Crossbars
- Improving Branch Prediction By Modeling Global History with Convolutional Neural Networks
- Knowledge Squeezed Adversarial Network Compression
- PruneNet: Channel Pruning via Global Importance
- Precision Highway for Ultra Low-Precision Quantization
- Channel-wise Hessian Aware trace-Weighted Quantization of Neural Networks
- Hardware-Guided Symbiotic Training for Compact, Accurate, yet Execution-Efficient LSTM
- Incremental Learning Using a Grow-and-Prune Paradigm with Efficient Neural Networks
- GAN Slimming: All-in-One GAN Compression by A Unified Optimization Framework
- Fractional Skipping: Towards Finer-Grained Dynamic CNN Inference
- Towards Lossless Binary Convolutional Neural Networks Using Piecewise Approximation
- Term Revealing: Furthering Quantization at Run Time on Quantized DNNs
- Ternary Residual Networks
- WrapNet: Neural Net Inference with Ultra-Low-Resolution Arithmetic
- Stochastic Downsampling for Cost-Adjustable Inference and Improved Regularization in Convolutional Networks
- Binarizing MobileNet via Evolution-based Searching
- Phoenix: A Low-Precision Floating-Point Quantization Oriented Architecture for Convolutional Neural Networks
- SparCE: Sparsity aware General Purpose Core Extensions to Accelerate Deep Neural Networks
- How Not to Give a FLOP: Combining Regularization and Pruning for Efficient Inference
- Search What You Want: Barrier Panelty NAS for Mixed Precision Quantization
- Distributed Low Precision Training Without Mixed Precision
- Sampling-Free Learning of Bayesian Quantized Neural Networks
- DiabDeep: Pervasive Diabetes Diagnosis based on Wearable Medical Sensors and Efficient Neural Networks
- An FPGA Accelerated Method for Training Feed-forward Neural Networks Using Alternating Direction Method of Multipliers and LSMR
- NASB: Neural Architecture Search for Binary Convolutional Neural Networks
- Learning Multimodal Fixed-Point Weights using Gradient Descent
- RTN: Reparameterized Ternary Network
- Deep Neural Network Approximation using Tensor Sketching
- SGQuant: Squeezing the Last Bit on Graph Neural Networks with Specialized Quantization
- WaveQ: Gradient-Based Deep Quantization of Neural Networks through Sinusoidal Adaptive Regularization
- Efficient Inferencing of Compressed Deep Neural Networks
- Faster Secure Data Mining via Distributed Homomorphic Encryption
- Information-Theoretic Understanding of Population Risk Improvement with Model Compression
- A Main/Subsidiary Network Framework for Simplifying Binary Neural Network
- Quantized Neural Network Inference with Precision Batching
- EDEN: Enabling Energy-Efficient, High-Performance Deep Neural Network Inference Using Approximate DRAM
- Towards thinner convolutional neural networks through Gradually Global Pruning
- Entropy-Based Modeling for Estimating Soft Errors Impact on Binarized Neural Network Inference
- -LBI: Stochastic Split Linearized Bregman Iterations for Parsimonious Deep Learning
- Self-grouping Convolutional Neural Networks
- Cluster Regularized Quantization for Deep Networks Compression
- A Targeted Acceleration and Compression Framework for Low bit Neural Networks
- Full-stack Optimization for Accelerating CNNs with FPGA Validation
- Cross-filter compression for CNN inference acceleration