XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks
arXiv:1603.05279
Abstract
We propose two efficient approximations to standard convolutional neural networks: Binary-Weight-Networks and XNOR-Networks. In Binary-Weight-Networks, the filters are approximated with binary values resulting in 32x memory saving. In XNOR-Networks, both the filters and the input to convolutional layers are binary. XNOR-Networks approximate convolutions using primarily binary operations. This results in 58x faster convolutional operations and 32x memory savings. XNOR-Nets offer the possibility of running state-of-the-art networks on CPUs (rather than GPUs) in real-time. Our binary networks are simple, accurate, efficient, and work on challenging visual tasks. We evaluate our approach on the ImageNet classification task. The classification accuracy with a Binary-Weight-Network version of AlexNet is only 2.9% less than the full-precision AlexNet (in top-1 measure). We compare our method with recent network binarization methods, BinaryConnect and BinaryNets, and outperform these methods by large margins on ImageNet, more than 16% in top-1 accuracy.
References in corpus (8)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and <0.5MB model size
- Deep Residual Learning for Image Recognition
- Compressing Deep Convolutional Networks using Vector Quantization
- Compressing Neural Networks with the Hashing Trick
- Speeding up Convolutional Neural Networks with Low Rank Expansions
- Bitwise Neural Networks
- Big Neural Networks Waste Capacity
Cited by in corpus (82)
- MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
- DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients
- Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations
- Ternary Weight Networks
- Trained Ternary Quantization
- Pruning Filters for Efficient ConvNets
- Demystifying Parallel and Distributed Deep Learning: An In-Depth Concurrency Analysis
- What-and-Where to Match: Deep Spatially Multiplicative Integration Networks for Person Re-identification
- DeepCABAC: A Universal Compression Algorithm for Deep Neural Networks
- Ristretto: Hardware-Oriented Approximation of Convolutional Neural Networks
- Deep Convolutional Neural Network Inference with Floating-point Weights and Fixed-point Activations
- Embedded Binarized Neural Networks
- Deep neural networks are robust to weight binarization and other non-linear distortions
- Distributed Deep Neural Networks over the Cloud, the Edge and End Devices
- NSGA-Net: Neural Architecture Search using Multi-Objective Genetic Algorithm
- Effective Quantization Methods for Recurrent Neural Networks
- SLSNet: Skin lesion segmentation using a lightweight generative adversarial network
- A neural network memory prefetcher using semantic locality
- Towards the Limit of Network Quantization
- Trained Quantization Thresholds for Accurate and Efficient Fixed-Point Inference of Deep Neural Networks
- Compacting Deep Neural Networks for Internet of Things: Methods and Applications
- ShiftCNN: Generalized Low-Precision Architecture for Inference of Convolutional Neural Networks
- A Survey on Deep Learning Methods for Robot Vision
- Delta Networks for Optimized Recurrent Network Computation
- Searching for Low-Bit Weights in Quantized Neural Networks
- Machine Learning Models that Remember Too Much
- FINN-R: An End-to-End Deep-Learning Framework for Fast Exploration of Quantized Neural Networks
- Streamlined Deployment for Quantized Neural Networks
- Training DNNs with Hybrid Block Floating Point
- Exploration of Low Numeric Precision Deep Learning Inference Using Intel FPGAs
- Not All Ops Are Created Equal!
- ShaResNet: reducing residual network parameter number by sharing weights
- Binarized Neural Networks on the ImageNet Classification Task
- Accelerating Deep Convolutional Networks using low-precision and sparsity
- Quantization and Training of Low Bit-Width Convolutional Neural Networks for Object Detection
- An Experimental Study of the Impact of Pre-training on the Pruning of a Convolutional Neural Network
- A Configurable BNN ASIC using a Network of Programmable Threshold Logic Standard Cells
- Local Feature Detectors, Descriptors, and Image Representations: A Survey
- Towards Evolutional Compression
- Toolflows for Mapping Convolutional Neural Networks on FPGAs: A Survey and Future Directions
- PAMS: Quantized Super-Resolution via Parameterized Max Scale
- Efficient Stochastic Inference of Bitwise Deep Neural Networks
- Recent Advances in the Applications of Convolutional Neural Networks to Medical Image Contour Detection
- Structured Convolution Matrices for Energy-efficient Deep learning
- Full deep neural network training on a pruned weight budget
- BinaryRelax: A Relaxation Approach For Training Deep Neural Networks With Quantized Weights
- HEMP: High-order Entropy Minimization for neural network comPression
- Training Bit Fully Convolutional Network for Fast Semantic Segmentation
- Optimizing for Interpretability in Deep Neural Networks with Tree Regularization
- Structured Deep Neural Network Pruning via Matrix Pivoting
- Efficient Hybrid Network Architectures for Extremely Quantized Neural Networks Enabling Intelligence at the Edge
- Accelerating Binarized Neural Networks via Bit-Tensor-Cores in Turing GPUs
- Fractional Skipping: Towards Finer-Grained Dynamic CNN Inference
- Automating Generation of Low Precision Deep Learning Operators
- Convolutional Neural Network Simplification with Progressive Retraining
- HarDNet: A Low Memory Traffic Network
- CNN Acceleration by Low-rank Approximation with Quantized Factors
- Low-Power Computer Vision: Status, Challenges, Opportunities
- A GPU-Outperforming FPGA Accelerator Architecture for Binary Convolutional Neural Networks
- Binarized Convolutional Neural Networks with Separable Filters for Efficient Hardware Acceleration
- A High-Performance Adaptive Quantization Approach for Edge CNN Applications
- DeepLight: Deep Lightweight Feature Interactions for Accelerating CTR Predictions in Ad Serving
- The Challenge of Multi-Operand Adders in CNNs on FPGAs: How not to solve it!
- Towards Fast and Energy-Efficient Binarized Neural Network Inference on FPGA
- Fast Video Classification via Adaptive Cascading of Deep Models
- On Psychoacoustically Weighted Cost Functions Towards Resource-Efficient Deep Neural Networks for Speech Denoising
- Always-On 674uW @ 4GOP/s Error Resilient Binary Neural Networks with Aggressive SRAM Voltage Scaling on a 22nm IoT End-Node
- DNN Feature Map Compression using Learned Representation over GF(2)
- Understanding the Energy and Precision Requirements for Online Learning
- Quantized neural network design under weight capacity constraint
- Neural network compression via learnable wavelet transforms
- Coarse and fine-grained automatic cropping deep convolutional neural network
- WeClick: Weakly-Supervised Video Semantic Segmentation with Click Annotations
- In-memory Implementation of On-chip Trainable and Scalable ANN for AI/ML Applications
- Scientific Calculator for Designing Trojan Detectors in Neural Networks
- Multi-Precision Quantized Neural Networks via Encoding Decomposition of -1 and +1
- Auto Deep Compression by Reinforcement Learning Based Actor-Critic Structure
- AE-Netv2: Optimization of Image Fusion Efficiency and Network Architecture
- clcNet: Improving the Efficiency of Convolutional Neural Network using Channel Local Convolutions
- Reliable Identification of Redundant Kernels for Convolutional Neural Network Compression
- Prune the Convolutional Neural Networks with Sparse Shrink
- Modulated binary cliquenet