BinaryConnect: Training Deep Neural Networks with binary weights during propagations
arXiv:1511.00363
Abstract
Deep Neural Networks (DNN) have achieved state-of-the-art results in a wide range of tasks, with the best results obtained with large training sets and large models. In the past, GPUs enabled these breakthroughs because of their greater computational speed. In the future, faster computation at both training and test time is likely to be crucial for further progress and for consumer applications on low-power devices. As a result, there is much interest in research and development of dedicated hardware for Deep Learning (DL). Binary weights, i.e., weights which are constrained to only two possible values (e.g. -1 or 1), would bring great benefits to specialized DL hardware by replacing many multiply-accumulate operations by simple accumulations, as multipliers are the most space and power-hungry components of the digital implementation of neural networks. We introduce BinaryConnect, a method which consists in training a DNN with binary weights during the forward and backward propagations, while retaining precision of the stored weights in which gradients are accumulated. Like other dropout schemes, we show that BinaryConnect acts as regularizer and we obtain near state-of-the-art results with BinaryConnect on the permutation-invariant MNIST, CIFAR-10 and SVHN.
Accepted at NIPS 2015, 9 pages, 3 figures
References in corpus (7)
- Sequence to Sequence Learning with Neural Networks
- Expectation Propagation for approximate Bayesian inference
- Going Deeper with Convolutions
- Theano: new features and speed improvements
- Spatially-sparse convolutional neural networks
- Rounding Methods for Neural Networks with Low Resolution Synaptic Weights
- Training Binary Multilayer Neural Networks for Image Classification using Expectation Backpropagation
Cited by in corpus (279)
- Knowledge Distillation: A Survey
- Binarized Neural Networks: Training Deep Neural Networks with Weights and Activations Constrained to +1 or -1
- AutoML: A Survey of the State-of-the-Art
- Convergence of Edge Computing and Deep Learning: A Comprehensive Survey
- Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations
- Deep Face Recognition: A Survey
- Mixed Precision Training
- A Survey of Model Compression and Acceleration for Deep Neural Networks
- Convolutional Networks for Fast, Energy-Efficient Neuromorphic Computing
- Ternary Weight Networks
- Trained Ternary Quantization
- Group Sparse Regularization for Deep Neural Networks
- Optimization Problems for Machine Learning: A Survey
- Soft-to-Hard Vector Quantization for End-to-End Learning Compressible Representations
- Bayesian Compression for Deep Learning
- Training and Inference with Integers in Deep Neural Networks
- NullHop: A Flexible Convolutional Neural Network Accelerator Based on Sparse Representations of Feature Maps
- XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks
- Hardware-oriented Approximation of Convolutional Neural Networks
- Model compression via distillation and quantization
- Hierarchical Multiscale Recurrent Neural Networks
- A Comprehensive Survey on Model Quantization for Deep Neural Networks in Image Classification
- Hardware and Software Optimizations for Accelerating Deep Neural Networks: Survey of Current Trends, Challenges, and the Road Ahead
- First-spike based visual categorization using reward-modulated STDP
- Demystifying Parallel and Distributed Deep Learning: An In-Depth Concurrency Analysis
- A Survey on Methods and Theories of Quantized Neural Networks
- WRPN: Wide Reduced-Precision Networks
- Online Batch Selection for Faster Training of Neural Networks
- Zoneout: Regularizing RNNs by Randomly Preserving Hidden Activations
- Apprentice: Using Knowledge Distillation Techniques To Improve Low-Precision Network Accuracy
- Distilled Siamese Networks for Visual Tracking
- Unreasonable Effectiveness of Learning Neural Networks: From Accessible States and Robust Ensembles to Basic Algorithmic Schemes
- An Entropy-based Pruning Method for CNN Compression
- Resiliency of Deep Neural Networks under Quantization
- Dynamic Network Surgery for Efficient DNNs
- Enable Deep Learning on Mobile Devices: Methods, Systems, and Applications
- Hardware Approximate Techniques for Deep Neural Network Accelerators: A Survey
- XNOR Neural Engine: a Hardware Accelerator IP for 21.6 fJ/op Binary Neural Network Inference
- TVM: An Automated End-to-End Optimizing Compiler for Deep Learning
- Recurrent Neural Networks: An Embedded Computing Perspective
- Mixed Precision Training of Convolutional Neural Networks using Integer Operations
- Compressing Neural Networks using the Variational Information Bottleneck
- Mixed-precision deep learning based on computational memory
- Knowledge Distillation in Deep Learning and its Applications
- Variable Rate Image Compression with Recurrent Neural Networks
- Training with Quantization Noise for Extreme Model Compression
- Compressing Recurrent Neural Network with Tensor Train
- Accelerating CNN inference on FPGAs: A Survey
- Ristretto: Hardware-Oriented Approximation of Convolutional Neural Networks
- LeViT: a Vision Transformer in ConvNet's Clothing for Faster Inference
- The Power of Sparsity in Convolutional Neural Networks
- Loss-aware Weight Quantization of Deep Networks
- BS4NN: Binarized Spiking Neural Networks with Temporal Coding and Learning
- NeuralPower: Predict and Deploy Energy-Efficient Convolutional Neural Networks
- Deep Neural Network Approximation for Custom Hardware: Where We've Been, Where We're Going
- Autoencoders on FPGAs for real-time, unsupervised new physics detection at 40 MHz at the Large Hadron Collider
- Deep neural networks are robust to weight binarization and other non-linear distortions
- Secure Evaluation of Quantized Neural Networks
- Embedded Binarized Neural Networks
- Distributed Deep Neural Networks over the Cloud, the Edge and End Devices
- AddNet: Deep Neural Networks Using FPGA-Optimized Multipliers
- Deep Learning with Low Precision by Half-wave Gaussian Quantization
- Continual Learning in Sensor-based Human Activity Recognition: an Empirical Benchmark Analysis
- And the Bit Goes Down: Revisiting the Quantization of Neural Networks
- MaskConnect: Connectivity Learning by Gradient Descent
- TAPAS: Tricks to Accelerate (encrypted) Prediction As a Service
- Compact recurrent neural networks for acoustic event detection on low-energy low-complexity platforms
- TBD: Benchmarking and Analyzing Deep Neural Network Training
- Automated Pruning for Deep Neural Network Compression
- Deep Spiking Networks
- Automated Verification of Neural Networks: Advances, Challenges and Perspectives
- Low-complexity Approximate Convolutional Neural Networks
- DFTerNet: Towards 2-bit Dynamic Fusion Networks for Accurate Human Activity Recognition
- An Optimal Control Approach to Deep Learning and Applications to Discrete-Weight Neural Networks
- HRel: Filter Pruning based on High Relevance between Activation Maps and Class Labels
- Efficient Sparse-Winograd Convolutional Neural Networks
- BinaryBERT: Pushing the Limit of BERT Quantization
- A Survey on Symbolic Knowledge Distillation of Large Language Models
- Delta Networks for Optimized Recurrent Network Computation
- Hierarchical binary CNNs for landmark localization with limited resources
- RANC: Reconfigurable Architecture for Neuromorphic Computing
- More is Less: A More Complicated Network with Less Inference Complexity
- Progressive Weight Pruning of Deep Neural Networks using ADMM
- Quantization of Fully Convolutional Networks for Accurate Biomedical Image Segmentation
- Compact Deep Convolutional Neural Networks With Coarse Pruning
- On-device Training: A First Overview on Existing Systems
- An Improvement of Data Classification Using Random Multimodel Deep Learning (RMDL)
- Piggyback: Adapting a Single Network to Multiple Tasks by Learning to Mask Weights
- Sparsely-Connected Neural Networks: Towards Efficient VLSI Implementation of Deep Neural Networks
- A Learning Framework for n-bit Quantized Neural Networks toward FPGAs
- ExPAN(N)D: Exploring Posits for Efficient Artificial Neural Network Design in FPGA-based Systems
- Controlling Information Capacity of Binary Neural Network
- Energy Efficient Learning with Low Resolution Stochastic Domain Wall Synapse Based Deep Neural Networks
- Rethinking Weight Decay For Efficient Neural Network Pruning
- Distributed Learning in Wireless Networks: Recent Progress and Future Challenges
- Learning Sparse Low-Precision Neural Networks With Learnable Regularization
- Model compression as constrained optimization, with application to neural nets. Part I: general framework
- Robustness of Neural Networks against Storage Media Errors
- How fine can fine-tuning be? Learning efficient language models
- Training DNNs with Hybrid Block Floating Point
- Communication-Efficient Edge AI: Algorithms and Systems
- Learning Accurate Low-Bit Deep Neural Networks with Stochastic Quantization
- Computing Systems for Autonomous Driving: State-of-the-Art and Challenges
- Connectivity Learning in Multi-Branch Networks
- Exploration of Low Numeric Precision Deep Learning Inference Using Intel FPGAs
- Quantization for Rapid Deployment of Deep Neural Networks
- Enhancing Neural Architecture Search with Multiple Hardware Constraints for Deep Learning Model Deployment on Tiny IoT Devices
- Programming and Training Rate-Independent Chemical Reaction Networks
- Model compression as constrained optimization, with application to neural nets. Part II: quantization
- AdaptivFloat: A Floating-point based Data Type for Resilient Deep Learning Inference
- Towards Effective Low-bitwidth Convolutional Neural Networks
- Optimally Scheduling CNN Convolutions for Efficient Memory Access
- Q-CapsNets: A Specialized Framework for Quantizing Capsule Networks
- Towards Unified INT8 Training for Convolutional Neural Network
- Computer Vision Model Compression Techniques for Embedded Systems: A Survey
- Adaptive Low-Precision Training for Embeddings in Click-Through Rate Prediction
- Deep Learning for Real-Time Crime Forecasting and its Ternarization
- Regularizing Activation Distribution for Training Binarized Deep Networks
- Degree-Quant: Quantization-Aware Training for Graph Neural Networks
- Differentiable Model Compression via Pseudo Quantization Noise
- Performance Guaranteed Network Acceleration via High-Order Residual Quantization
- Accuracy-Efficiency Trade-Offs and Accountability in Distributed ML Systems
- Neural Networks Designing Neural Networks: Multi-Objective Hyper-Parameter Optimization
- Tensorial Neural Networks: Generalization of Neural Networks and Application to Model Compression
- Binarized Neural Networks on the ImageNet Classification Task
- Projection Convolutional Neural Networks for 1-bit CNNs via Discrete Back Propagation
- Efficient Visual Recognition with Deep Neural Networks: A Survey on Recent Advances and New Directions
- Mixed Precision Low-bit Quantization of Neural Network Language Models for Speech Recognition
- Incremental Binarization On Recurrent Neural Networks For Single-Channel Source Separation
- Reverse Derivative Ascent: A Categorical Approach to Learning Boolean Circuits
- Taxonomy and Benchmarking of Precision-Scalable MAC Arrays Under Enhanced DNN Dataflow Representation
- Dynamic Hard Pruning of Neural Networks at the Edge of the Internet
- Iteratively Training Look-Up Tables for Network Quantization
- Structured Binary Neural Networks for Accurate Image Classification and Semantic Segmentation
- Tartan: Accelerating Fully-Connected and Convolutional Layers in Deep Learning Networks by Exploiting Numerical Precision Variability
- Fine-Pruning: Joint Fine-Tuning and Compression of a Convolutional Network with Bayesian Optimization
- Quantization and Training of Low Bit-Width Convolutional Neural Networks for Object Detection
- BitPruning: Learning Bitlengths for Aggressive and Accurate Quantization
- Binary Ensemble Neural Network: More Bits per Network or More Networks per Bit?
- An Encoding Framework for Binarized Images using HyperDimensional Computing
- Exploiting Errors for Efficiency: A Survey from Circuits to Algorithms
- Searching for Winograd-aware Quantized Networks
- In-situ Stochastic Training of MTJ Crossbar based Neural Networks
- SEP-Nets: Small and Effective Pattern Networks
- On Periodic Functions as Regularizers for Quantization of Neural Networks
- OpenEI: An Open Framework for Edge Intelligence
- DaCapo: Accelerating Continuous Learning in Autonomous Systems for Video Analytics
- Fix your classifier: the marginal value of training the last weight layer
- LCNN: Lookup-based Convolutional Neural Network
- Towards Evolutional Compression
- Low-Precision Batch-Normalized Activations
- Efficient Stochastic Inference of Bitwise Deep Neural Networks
- Adjustable Bounded Rectifiers: Towards Deep Binary Representations
- BiFSMN: Binary Neural Network for Keyword Spotting
- Mixed-precision training of deep neural networks using computational memory
- Temporally Efficient Deep Learning with Spikes
- Joint Pruning & Quantization for Extremely Sparse Neural Networks
- Sparse Weight Activation Training
- BMXNet: An Open-Source Binary Neural Network Implementation Based on MXNet
- Espresso: Efficient Forward Propagation for BCNNs
- Depthwise Multiception Convolution for Reducing Network Parameters without Sacrificing Accuracy
- HEMP: High-order Entropy Minimization for neural network comPression
- Channel-wise Hessian Aware trace-Weighted Quantization of Neural Networks
- Deep Feature Flow for Video Recognition
- SYMOG: learning symmetric mixture of Gaussian modes for improved fixed-point quantization
- Learning Versatile Convolution Filters for Efficient Visual Recognition
- Accelerating Deterministic and Stochastic Binarized Neural Networks on FPGAs Using OpenCL
- Deep causal representation learning for unsupervised domain adaptation
- Propagating Asymptotic-Estimated Gradients for Low Bitwidth Quantized Neural Networks
- Understanding the Impact of Precision Quantization on the Accuracy and Energy of Neural Networks
- Divide and Conquer: Leveraging Intermediate Feature Representations for Quantized Training of Neural Networks
- GXNOR-Net: Training deep neural networks with ternary weights and activations without full-precision memory under a unified discretization framework
- Tag Prediction at Flickr: a View from the Darkroom
- Recursive Binary Neural Network Learning Model with 2.28b/Weight Storage Requirement
- Deep Dose Plugin Towards Real-time Monte Carlo Dose Calculation Through a Deep Learning based Denoising Algorithm
- Nonuniform-to-Uniform Quantization: Towards Accurate Quantization via Generalized Straight-Through Estimation
- Getting deep recommenders fit: Bloom embeddings for sparse binary input/output networks
- On the efficient representation and execution of deep acoustic models
- At-Scale Evaluation of Weight Clustering to Enable Energy-Efficient Object Detection
- Taming Binarized Neural Networks and Mixed-Integer Programs
- SparCE: Sparsity aware General Purpose Core Extensions to Accelerate Deep Neural Networks
- CNN Acceleration by Low-rank Approximation with Quantized Factors
- Multi-Complexity-Loss DNAS for Energy-Efficient and Memory-Constrained Deep Neural Networks
- DEEPEYE: A Compact and Accurate Video Comprehension at Terminal Devices Compressed with Quantization and Tensorization
- Efficient Hardware Realization of Convolutional Neural Networks using Intra-Kernel Regular Pruning
- BitNet: Bit-Regularized Deep Neural Networks
- Compressing complex convolutional neural network based on an improved deep compression algorithm
- Quantized Non-Volatile Nanomagnetic Synapse based Autoencoder for Efficient Unsupervised Network Anomaly Detection
- A Bop and Beyond: A Second Order Optimizer for Binarized Neural Networks
- A GPU-Outperforming FPGA Accelerator Architecture for Binary Convolutional Neural Networks
- Distilled Non-Semantic Speech Embeddings with Binary Neural Networks for Low-Resource Devices
- A Simple Method to Reduce Off-chip Memory Accesses on Convolutional Neural Networks
- Differentiable Sparsification for Deep Neural Networks
- Binarized Convolutional Neural Networks with Separable Filters for Efficient Hardware Acceleration
- Activation Compression of Graph Neural Networks using Block-wise Quantization with Improved Variance Minimization
- Training for 'Unstable' CNN Accelerator:A Case Study on FPGA
- Knowledge distillation for optimization of quantized deep neural networks
- Generalized Ternary Connect: End-to-End Learning and Compression of Multiplication-Free Deep Neural Networks
- Approximations in Deep Learning
- Convolutional Neural Network Quantization using Generalized Gamma Distribution
- Training Deep Neural Network in Limited Precision
- A High-Performance Adaptive Quantization Approach for Edge CNN Applications
- Binary Complex Neural Network Acceleration on FPGA
- Matrix and tensor decompositions for training binary neural networks
- Detecting Dead Weights and Units in Neural Networks
- Privacy-Preserving Visual Learning Using Doubly Permuted Homomorphic Encryption
- Simplified Stochastic Feedforward Neural Networks
- Some Remarks on Replicated Simulated Annealing
- Digital Neuron: A Hardware Inference Accelerator for Convolutional Deep Neural Networks
- Compressing LSTM Networks by Matrix Product Operators
- Recent Advances in Efficient Computation of Deep Convolutional Neural Networks
- Embedded Knowledge Distillation in Depth-Level Dynamic Neural Network
- A Selective Survey on Versatile Knowledge Distillation Paradigm for Neural Network Models
- Hierarchical compositional feature learning
- Toward Computation and Memory Efficient Neural Network Acoustic Models with Binary Weights and Activations
- Learning Recurrent Binary/Ternary Weights
- Low Precision Policy Distillation with Application to Low-Power, Real-time Sensation-Cognition-Action Loop with Neuromorphic Computing
- Towards Fast and Energy-Efficient Binarized Neural Network Inference on FPGA
- n-hot: Efficient bit-level sparsity for powers-of-two neural network quantization
- TinBiNN: Tiny Binarized Neural Network Overlay in about 5,000 4-LUTs and 5mW
- AskewSGD : An Annealed interval-constrained Optimisation method to train Quantized Neural Networks
- Integer-Only Neural Network Quantization Scheme Based on Shift-Batch-Normalization
- Recent Advances in Convolutional Neural Network Acceleration
- Binary-decomposed DCNN for accelerating computation and compressing model without retraining
- Leveraging Structured Pruning of Convolutional Neural Networks
- BCNN: Binary Complex Neural Network
- Balanced Quantization: An Effective and Efficient Approach to Quantized Neural Networks
- Training of Quantized Deep Neural Networks using a Magnetic Tunnel Junction-Based Synapse
- Binarized Knowledge Graph Embeddings
- Fixed-point optimization of deep neural networks with adaptive step size retraining
- Learning Multimodal Fixed-Point Weights using Gradient Descent
- Diagonalwise Refactorization: An Efficient Training Method for Depthwise Convolutions
- Local Binary Pattern Networks
- Learning Instance-wise Sparsity for Accelerating Deep Models
- Energy Consumption Analysis of pruned Semantic Segmentation Networks on an Embedded GPU
- Blind Adversarial Pruning: Balance Accuracy, Efficiency and Robustness
- A Main/Subsidiary Network Framework for Simplifying Binary Neural Network
- Efficient and Robust Machine Learning for Real-World Systems
- Classification Accuracy Improvement for Neuromorphic Computing Systems with One-level Precision Synapses
- Composite Binary Decomposition Networks
- On Transformations in Stochastic Gradient MCMC
- Finding Everything within Random Binary Networks
- Quantization Mimic: Towards Very Tiny CNN for Object Detection
- NEAT: A Framework for Automated Exploration of Floating Point Approximations
- A New Training Framework for Deep Neural Network
- "BNN - BN = ?": Training Binary Neural Networks without Batch Normalization
- Convolutional neural networks compression with low rank and sparse tensor decompositions
- Markov chain Hebbian learning algorithm with ternary synaptic units
- Understanding the Energy and Precision Requirements for Online Learning
- A Layer Decomposition-Recomposition Framework for Neuron Pruning towards Accurate Lightweight Networks
- Quantized neural network design under weight capacity constraint
- Automated Deep Abstractions for Stochastic Chemical Reaction Networks
- Maximin Optimization for Binary Regression
- RMSMP: A Novel Deep Neural Network Quantization Framework with Row-wise Mixed Schemes and Multiple Precisions
- Selective Deep Convolutional Neural Network for Low Cost Distorted Image Classification
- Optimize Deep Convolutional Neural Network with Ternarized Weights and High Accuracy
- Binarized Convolutional Neural Networks for Efficient Inference on GPUs
- A study on speech enhancement using exponent-only floating point quantized neural network (EOFP-QNN)
- Robust Implicit Backpropagation
- An Overview of Datatype Quantization Techniques for Convolutional Neural Networks
- Fixed-point Factorized Networks
- ATP-Net: An Attention-based Ternary Projection Network For Compressed Sensing
- Training compact deep learning models for video classification using circulant matrices
- Designing Energy-Efficient Convolutional Neural Networks using Energy-Aware Pruning
- Penetrating the Fog: the Path to Efficient CNN Models
- ADA-Tucker: Compressing Deep Neural Networks via Adaptive Dimension Adjustment Tucker Decomposition
- A Greedy Algorithm to Cluster Specialists
- Learning to Skip Ineffectual Recurrent Computations in LSTMs
- Circulant Binary Convolutional Networks: Enhancing the Performance of 1-bit DCNNs with Circulant Back Propagation
- Reliable Identification of Redundant Kernels for Convolutional Neural Network Compression
- Multi-Precision Quantized Neural Networks via Encoding Decomposition of -1 and +1
- : Sign-Sparse-Shift Reparametrization for Effective Training of Low-bit Shift Networks
- Bayesian Optimized 1-Bit CNNs
- Adaptive Binary-Ternary Quantization
- Collaborative Deep Learning for Speech Enhancement: A Run-Time Model Selection Method Using Autoencoders
- Neural Network Activation Quantization with Bitwise Information Bottlenecks
- ENOS: Energy-Aware Network Operator Search for Hybrid Digital and Compute-in-Memory DNN Accelerators
- Homomorphic Parameter Compression for Distributed Deep Learning Training
- Recurrent Residual Module for Fast Inference in Videos