Ternary Weight Networks
arXiv:1605.04711
Abstract
We present a memory and computation efficient ternary weight networks (TWNs) - with weights constrained to +1, 0 and -1. The Euclidian distance between full (float or double) precision weights and the ternary weights along with a scaling factor is minimized in training stage. Besides, a threshold-based ternary function is optimized to get an approximated solution which can be fast and easily computed. TWNs have shown better expressive abilities than binary precision counterparts. Meanwhile, TWNs achieve up to 16 model compression rate and need fewer multiplications compared with the float32 precision counterparts. Extensive experiments on MNIST, CIFAR-10, and ImageNet datasets show that the TWNs achieve much better result than the Binary-Weight-Networks (BWNs) and the classification performance on MNIST and CIFAR-10 is very close to the full precision networks. We also verify our method on object detection task and show that TWNs significantly outperforms BWN by more than 10\% mAP on PASCAL VOC dataset. The pytorch version of source code is available at: https://github.com/Thinklab-SJTU/twns.
5 pages, 3 fitures, conference
References in corpus (4)
Cited by in corpus (248)
- DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients
- Trained Ternary Quantization
- Efficient Deep Learning: A Survey on Making Deep Learning Models Smaller, Faster, and Better
- Distillation-Based Semi-Supervised Federated Learning for Communication-Efficient Collaborative Training with Non-IID Private Data
- Training and Inference with Integers in Deep Neural Networks
- Learned Step Size Quantization
- Hardware and Software Optimizations for Accelerating Deep Neural Networks: Survey of Current Trends, Challenges, and the Road Ahead
- Data-Free Knowledge Distillation for Deep Neural Networks
- Demystifying Parallel and Distributed Deep Learning: An In-Depth Concurrency Analysis
- A Survey on Methods and Theories of Quantized Neural Networks
- WRPN: Wide Reduced-Precision Networks
- Bringing AI To Edge: From Deep Learning's Perspective
- Discrimination-aware Channel Pruning for Deep Neural Networks
- Apprentice: Using Knowledge Distillation Techniques To Improve Low-Precision Network Accuracy
- A Survey of FPGA-Based Neural Network Accelerator
- Enable Deep Learning on Mobile Devices: Methods, Systems, and Applications
- Recurrent Neural Networks: An Embedded Computing Perspective
- An End to End Deep Neural Network for Iris Segmentation in Unconstraint Scenarios
- Exploring the Connection Between Binary and Spiking Neural Networks
- DeepCABAC: A Universal Compression Algorithm for Deep Neural Networks
- Adaptive Inference through Early-Exit Networks: Design, Challenges and Directions
- Rethinking floating point for deep learning
- Accelerating CNN inference on FPGAs: A Survey
- Layer-specific Optimization for Mixed Data Flow with Mixed Precision in FPGA Design for CNN-based Object Detectors
- Training Quantized Nets: A Deeper Understanding
- Loss-aware Weight Quantization of Deep Networks
- SkyNet: a Hardware-Efficient Method for Object Detection and Tracking on Embedded Systems
- Big-Little Net: An Efficient Multi-Scale Feature Representation for Visual and Speech Recognition
- Bayesian Bits: Unifying Quantization and Pruning
- Deep Neural Network Approximation for Custom Hardware: Where We've Been, Where We're Going
- ProxQuant: Quantized Neural Networks via Proximal Operators
- Deep Learning with Low Precision by Half-wave Gaussian Quantization
- And the Bit Goes Down: Revisiting the Quantization of Neural Networks
- Ternary Neural Networks with Fine-Grained Quantization
- Understanding Straight-Through Estimator in Training Activation Quantized Neural Nets
- Differentiable Soft Quantization: Bridging Full-Precision and Low-Bit Neural Networks
- Trained Quantization Thresholds for Accurate and Efficient Fixed-Point Inference of Deep Neural Networks
- Compacting Deep Neural Networks for Internet of Things: Methods and Applications
- Crossbar-aware neural network pruning
- ShiftCNN: Generalized Low-Precision Architecture for Inference of Convolutional Neural Networks
- Q-BERT: Hessian Based Ultra Low Precision Quantization of BERT
- DFTerNet: Towards 2-bit Dynamic Fusion Networks for Accurate Human Activity Recognition
- Mixed Precision DNNs: All you need is a good parametrization
- An Optimal Control Approach to Deep Learning and Applications to Discrete-Weight Neural Networks
- QKD: Quantization-aware Knowledge Distillation
- Efficient Processing of Deep Neural Networks: A Tutorial and Survey
- BinaryBERT: Pushing the Limit of BERT Quantization
- TernaryNet: Faster Deep Model Inference without GPUs for Medical 3D Segmentation using Sparse and Binary Convolutions
- ReLeQ: A Reinforcement Learning Approach for Deep Quantization of Neural Networks
- Towards Unconstrained Palmprint Recognition on Consumer Devices: a Literature Review
- A Survey of FPGA Based Deep Learning Accelerators: Challenges and Opportunities
- Quantization of Fully Convolutional Networks for Accurate Biomedical Image Segmentation
- Bridging the Accuracy Gap for 2-bit Quantized Neural Networks (QNN)
- Piggyback: Adapting a Single Network to Multiple Tasks by Learning to Mask Weights
- A Learning Framework for n-bit Quantized Neural Networks toward FPGAs
- Towards Efficient Training for Neural Network Quantization
- DSConv: Efficient Convolution Operator
- Bit Fusion: Bit-Level Dynamically Composable Architecture for Accelerating Deep Neural Networks
- Musical Chair: Efficient Real-Time Recognition Using Collaborative IoT Devices
- ReActNet: Towards Precise Binary Neural Network with Generalized Activation Functions
- Binary Neural Networks for Memory-Efficient and Effective Visual Place Recognition in Changing Environments
- IGCV3: Interleaved Low-Rank Group Convolutions for Efficient Deep Neural Networks
- Training DNNs with Hybrid Block Floating Point
- Learning Accurate Low-Bit Deep Neural Networks with Stochastic Quantization
- Forward and Backward Information Retention for Accurate Binary Neural Networks
- Model compression as constrained optimization, with application to neural nets. Part I: general framework
- Quantization for Rapid Deployment of Deep Neural Networks
- Exploration of Low Numeric Precision Deep Learning Inference Using Intel FPGAs
- SquishedNets: Squishing SqueezeNet further for edge device scenarios via deep evolutionary synthesis
- Self-Binarizing Networks
- Mixed Low-precision Deep Learning Inference using Dynamic Fixed Point
- Model compression as constrained optimization, with application to neural nets. Part II: quantization
- Towards End-to-End Neural Face Authentication in the Wild -- Quantifying and Compensating for Directional Lighting Effects
- Towards Efficient Post-training Quantization of Pre-trained Language Models
- Resource-Efficient Neural Networks for Embedded Systems
- Deep Learning for Real-Time Crime Forecasting and its Ternarization
- ZeroQ: A Novel Zero Shot Quantization Framework
- Differentiable Model Compression via Pseudo Quantization Noise
- Collaborative Execution of Deep Neural Networks on Internet of Things Devices
- Compressive Sensing Using Iterative Hard Thresholding with Low Precision Data Representation: Theory and Applications
- AdaBits: Neural Network Quantization with Adaptive Bit-Widths
- Mixed Precision Low-bit Quantization of Neural Network Language Models for Speech Recognition
- Temporal-Coded Deep Spiking Neural Network with Easy Training and Robust Performance
- FQ-Conv: Fully Quantized Convolution for Efficient and Accurate Inference
- Intermediate Deep Feature Compression: the Next Battlefield of Intelligent Sensing
- Relaxed Quantization for Discretized Neural Networks
- Iteratively Training Look-Up Tables for Network Quantization
- Precision Gating: Improving Neural Network Efficiency with Dynamic Dual-Precision Activations
- I-BERT: Integer-only BERT Quantization
- Switchable Precision Neural Networks
- The ZipML Framework for Training Models with End-to-End Low Precision: The Cans, the Cannots, and a Little Bit of Deep Learning
- Quantization and Training of Low Bit-Width Convolutional Neural Networks for Object Detection
- Binary Ensemble Neural Network: More Bits per Network or More Networks per Bit?
- MoBiNet: A Mobile Binary Network for Image Classification
- Bringing Giant Neural Networks Down to Earth with Unlabeled Data
- In-situ Stochastic Training of MTJ Crossbar based Neural Networks
- Searching for Winograd-aware Quantized Networks
- Knowledge Squeezed Adversarial Network Compression
- Ternary Compression for Communication-Efficient Federated Learning
- On Periodic Functions as Regularizers for Quantization of Neural Networks
- Efficient Bitwidth Search for Practical Mixed Precision Neural Network
- Communication-Efficient Federated Distillation
- Fix your classifier: the marginal value of training the last weight layer
- Blended Coarse Gradient Descent for Full Quantization of Deep Neural Networks
- Quantized Densely Connected U-Nets for Efficient Landmark Localization
- A Lite Distributed Semantic Communication System for Internet of Things
- FleXOR: Trainable Fractional Quantization
- Bit Efficient Quantization for Deep Neural Networks
- Low-Precision Batch-Normalized Activations
- Lottery Hypothesis based Unsupervised Pre-training for Model Compression in Federated Learning
- CPT: Efficient Deep Neural Network Training via Cyclic Precision
- FxP-QNet: A Post-Training Quantizer for the Design of Mixed Low-Precision DNNs with Dynamic Fixed-Point Representation
- Eyeriss v2: A Flexible Accelerator for Emerging Deep Neural Networks on Mobile Devices
- TernaryBERT: Distillation-aware Ultra-low Bit BERT
- Weight Normalization based Quantization for Deep Neural Network Compression
- Sparse Weight Activation Training
- Joint Pruning & Quantization for Extremely Sparse Neural Networks
- Training Compact Neural Networks with Binary Weights and Low Precision Activations
- RPR: Random Partition Relaxation for Training; Binary and Ternary Weight Neural Networks
- BinaryRelax: A Relaxation Approach For Training Deep Neural Networks With Quantized Weights
- Structured Convolutions for Efficient Neural Network Design
- SYMOG: learning symmetric mixture of Gaussian modes for improved fixed-point quantization
- Network Quantization with Element-wise Gradient Scaling
- LUTNet: Rethinking Inference in FPGA Soft Logic
- Deep Molecular Programming: A Natural Implementation of Binary-Weight ReLU Neural Networks
- Scalable Model Compression by Entropy Penalized Reparameterization
- Simultaneously Optimizing Weight and Quantizer of Ternary Neural Network using Truncated Gaussian Approximation
- HadaNets: Flexible Quantization Strategies for Neural Networks
- MWQ: Multiscale Wavelet Quantized Neural Networks
- Propagating Asymptotic-Estimated Gradients for Low Bitwidth Quantized Neural Networks
- Learning Filter Basis for Convolutional Neural Network Compression
- Distillation Guided Residual Learning for Binary Convolutional Neural Networks
- Mix and Match: A Novel FPGA-Centric Deep Neural Network Quantization Framework
- Towards Lossless Binary Convolutional Neural Networks Using Piecewise Approximation
- Evolutionary Bin Packing for Memory-Efficient Dataflow Inference Acceleration on FPGA
- WrapNet: Neural Net Inference with Ultra-Low-Resolution Arithmetic
- Deep Dose Plugin Towards Real-time Monte Carlo Dose Calculation Through a Deep Learning based Denoising Algorithm
- Ternary Residual Networks
- GXNOR-Net: Training deep neural networks with ternary weights and activations without full-precision memory under a unified discretization framework
- LegoNet: Memory Footprint Reduction Through Block Weight Clustering
- SOTERIA: In Search of Efficient Neural Networks for Private Inference
- End-to-End Learned Image Compression with Quantized Weights and Activations
- Model compression as constrained optimization, with application to neural nets. Part V: combining compressions
- 3DQ: Compact Quantized Neural Networks for Volumetric Whole Brain Segmentation
- SinReQ: Generalized Sinusoidal Regularization for Low-Bitwidth Deep Quantized Training
- Design Flow of Accelerating Hybrid Extremely Low Bit-width Neural Network in Embedded FPGA
- A 1Mb mixed-precision quantized encoder for image classification and patch-based compression
- Learning Architectures for Binary Networks
- Dynamic Runtime Feature Map Pruning
- Differentiable Joint Pruning and Quantization for Hardware Efficiency
- Automatic Neural Network Compression by Sparsity-Quantization Joint Learning: A Constrained Optimization-based Approach
- Generative Zero-shot Network Quantization
- Quantization of Deep Neural Networks for Accurate Edge Computing
- MSP: An FPGA-Specific Mixed-Scheme, Multi-Precision Deep Neural Network Quantization Framework
- Optimal Gradient Quantization Condition for Communication-Efficient Distributed Training
- LCP: A Low-Communication Parallelization Method for Fast Neural Network Inference in Image Recognition
- A Unified DNN Weight Compression Framework Using Reweighted Optimization Methods
- A High-Performance Adaptive Quantization Approach for Edge CNN Applications
- Binarized Neural Architecture Search for Efficient Object Recognition
- Training Deep Neural Network in Limited Precision
- NASB: Neural Architecture Search for Binary Convolutional Neural Networks
- Binary Input Layer: Training of CNN models with binary input data
- DNQ: Dynamic Network Quantization
- Learning Sparse & Ternary Neural Networks with Entropy-Constrained Trained Ternarization (EC2T)
- Learning Recurrent Binary/Ternary Weights
- An FPGA Accelerated Method for Training Feed-forward Neural Networks Using Alternating Direction Method of Multipliers and LSMR
- Joint Architecture and Knowledge Distillation in CNN for Chinese Text Recognition
- Distribution-sensitive Information Retention for Accurate Binary Neural Network
- Automated flow for compressing convolution neural networks for efficient edge-computation with FPGA
- Time-multiplexed In-memory computation scheme for mapping Quantized Neural Networks on hybrid CMOS-OxRAM building blocks
- Efficient Integer-Arithmetic-Only Convolutional Neural Networks
- Recent Advances in Efficient Computation of Deep Convolutional Neural Networks
- Histogram-Equalized Quantization for logic-gated Residual Neural Networks
- Progressive Learning of Low-Precision Networks
- n-hot: Efficient bit-level sparsity for powers-of-two neural network quantization
- Deep Neural Network Approximation using Tensor Sketching
- BitHEP -- The Limits of Low-Precision ML in HEP
- Training of Quantized Deep Neural Networks using a Magnetic Tunnel Junction-Based Synapse
- BNAS v2: Learning Architectures for Binary Networks with Empirical Improvements
- Differentiable Neural Architecture Learning for Efficient Neural Network Design
- RTN: Reparameterized Ternary Network
- BILLNET: A Binarized Conv3D-LSTM Network with Logic-gated residual architecture for hardware-efficient video inference
- Structured Compression by Weight Encryption for Unstructured Pruning and Quantization
- Neural Networks Weights Quantization: Target None-retraining Ternary (TNT)
- Localization-aware Channel Pruning for Object Detection
- Running Neural Networks on the NIC
- A Survey of FPGA-Based Robotic Computing
- Recent Advances in Convolutional Neural Network Acceleration
- Reward-Based 1-bit Compressed Federated Distillation on Blockchain
- FATNN: Fast and Accurate Ternary Neural Networks
- Mitigate Parasitic Resistance in Resistive Crossbar-based Convolutional Neural Networks
- Training of mixed-signal optical convolutional neural network with reduced quantization level
- BiDet: An Efficient Binarized Object Detector
- Dual Precision Deep Neural Network
- Improving Network Slimming with Nonconvex Regularization
- LANCE: Efficient Low-Precision Quantized Winograd Convolution for Neural Networks Based on Graphics Processing Units
- Improved Techniques for Quantizing Deep Networks with Adaptive Bit-Widths
- 3U-EdgeAI: Ultra-Low Memory Training, Ultra-Low BitwidthQuantization, and Ultra-Low Latency Acceleration
- SQWA: Stochastic Quantized Weight Averaging for Improving the Generalization Capability of Low-Precision Deep Neural Networks
- Pruning Ternary Quantization
- Prune Your Model Before Distill It
- PSDNet and DPDNet: Efficient channel expansion, Depthwise-Pointwise-Depthwise Inverted Bottleneck Block
- MOGNET: A Mux-residual quantized Network leveraging Online-Generated weights
- Towards thinner convolutional neural networks through Gradually Global Pruning
- Efficient and Robust Machine Learning for Real-World Systems
- Is In-Domain Data Really Needed? A Pilot Study on Cross-Domain Calibration for Network Quantization
- Efficient Micro-Structured Weight Unification and Pruning for Neural Network Compression
- ILMPQ : An Intra-Layer Multi-Precision Deep Neural Network Quantization framework for FPGA
- Reducing Inference Latency with Concurrent Architectures for Image Recognition
- Low-memory convolutional neural networks through incremental depth-first processing
- Entropy-Based Modeling for Estimating Soft Errors Impact on Binarized Neural Network Inference
- PBGen: Partial Binarization of Deconvolution-Based Generators for Edge Intelligence
- End-to-end Learned Image Compression with Fixed Point Weight Quantization
- Lightweight Neural Networks
- Binarized Weight Error Networks With a Transition Regularization Term
- Optimize Deep Convolutional Neural Network with Ternarized Weights and High Accuracy
- DecisiveNets: Training Deep Associative Memories to Solve Complex Machine Learning Problems
- Adaptive Precision Training for Resource Constrained Devices
- Direct Quantization for Training Highly Accurate Low Bit-width Deep Neural Networks
- Quantized Neural Networks via {-1, +1} Encoding Decomposition and Acceleration
- Tetris: Re-architecting Convolutional Neural Network Computation for Machine Learning Accelerators
- AutoQNN: An End-to-End Framework for Automatically Quantizing Neural Networks
- DNN Quantization with Attention
- Faster Convolution Inference Through Using Pre-Calculated Lookup Tables
- Online Filter Clustering and Pruning for Efficient Convnets
- Sparsity-Control Ternary Weight Networks
- Self-grouping Convolutional Neural Networks
- Automatic low-bit hybrid quantization of neural networks through meta learning
- Cross-filter compression for CNN inference acceleration
- Low-bit Quantization of Recurrent Neural Network Language Models Using Alternating Direction Methods of Multipliers
- Mixed Precision DNN Qunatization for Overlapped Speech Separation and Recognition
- BitSplit-Net: Multi-bit Deep Neural Network with Bitwise Activation Function
- Memory and Computation-Efficient Kernel SVM via Binary Embedding and Ternary Model Coefficients
- Learning Quantized Neural Nets by Coarse Gradient Method for Non-linear Classification
- HOTCAKE: Higher Order Tucker Articulated Kernels for Deeper CNN Compression
- Resource-Efficient Speech Mask Estimation for Multi-Channel Speech Enhancement
- Cluster Regularized Quantization for Deep Networks Compression
- Toward Compact Parameter Representations for Architecture-Agnostic Neural Network Compression
- Iterative Training: Finding Binary Weight Deep Neural Networks with Layer Binarization
- Adaptive Binary-Ternary Quantization
- Semi-Relaxed Quantization with DropBits: Training Low-Bit Neural Networks via Bit-wise Regularization
- An Overview of Datatype Quantization Techniques for Convolutional Neural Networks
- CBP: Backpropagation with constraint on weight precision using a pseudo-Lagrange multiplier method
- Demystifying and Generalizing BinaryConnect
- : Sign-Sparse-Shift Reparametrization for Effective Training of Low-bit Shift Networks
- RMSMP: A Novel Deep Neural Network Quantization Framework with Row-wise Mixed Schemes and Multiple Precisions
- Towards Mixed-Precision Quantization of Neural Networks via Constrained Optimization
- Hardware realisation of nonlinear dynamical systems for and from biology