Variational Dropout Sparsifies Deep Neural Networks
arXiv:1701.05369
Abstract
We explore a recently proposed Variational Dropout technique that provided an elegant Bayesian interpretation to Gaussian Dropout. We extend Variational Dropout to the case when dropout rates are unbounded, propose a way to reduce the variance of the gradient estimator and report first experimental results with individual dropout rates per weight. Interestingly, it leads to extremely sparse solutions both in fully-connected and convolutional layers. This effect is similar to automatic relevance determination effect in empirical Bayes but has a number of advantages. We reduce the number of parameters up to 280 times on LeNet architectures and up to 68 times on VGG-like networks with a negligible decrease of accuracy.
Published in ICML 2017
References in corpus (6)
- Improving neural networks by preventing co-adaptation of feature detectors
- Understanding deep learning requires rethinking generalization
- Group Sparse Regularization for Deep Neural Networks
- Learning Structured Sparsity in Deep Neural Networks
- Ultimate tensorization: compressing convolutional and FC layers alike
- The Power of Sparsity in Convolutional Neural Networks
Cited by in corpus (155)
- An Introduction to Variational Autoencoders
- The State of Sparsity in Deep Neural Networks
- R-Drop: Regularized Dropout for Neural Networks
- Integration of Neural Network-Based Symbolic Regression in Deep Learning for Scientific Discovery
- Sparse Networks from Scratch: Faster Training without Losing Performance
- Rigging the Lottery: Making All Tickets Winners
- Training with Quantization Noise for Extreme Model Compression
- DeepCABAC: A Universal Compression Algorithm for Deep Neural Networks
- One ticket to win them all: generalizing lottery ticket initializations across datasets and optimizers
- Soft Threshold Weight Reparameterization for Learnable Sparsity
- What Are Bayesian Neural Network Posteriors Really Like?
- Playing the lottery with rewards and multiple languages: lottery tickets in RL and NLP
- Dynamic Model Pruning with Feedback
- SpArSe: Sparse Architecture Search for CNNs on Resource-Constrained Microcontrollers
- What is the State of Neural Network Pruning?
- Pitfalls of In-Domain Uncertainty Estimation and Ensembling in Deep Learning
- Layer-adaptive sparsity for the Magnitude-based Pruning
- Discovering Neural Wirings
- Survey of Dropout Methods for Deep Neural Networks
- Compression of Recurrent Neural Networks for Efficient Language Modeling
- MEST: Accurate and Fast Memory-Economic Sparse Training Framework on the Edge
- NICE: Noise Injection and Clamping Estimation for Neural Network Quantization
- Gradual Channel Pruning while Training using Feature Relevance Scores for Convolutional Neural Networks
- Efficient and Scalable Bayesian Neural Nets with Rank-1 Factors
- Information Aware Max-Norm Dirichlet Networks for Predictive Uncertainty Estimation
- LOss-Based SensiTivity rEgulaRization: towards deep sparse neural networks
- The Difficulty of Training Sparse Neural Networks
- Rethinking Weight Decay For Efficient Neural Network Pruning
- Uncertainty Quantification in Deep Learning for Safer Neuroimage Enhancement
- Learning Sparse Low-Precision Neural Networks With Learnable Regularization
- Continual Learning with Adaptive Weights (CLAW)
- Dynamic Sparse Training: Find Efficient Sparse Network From Scratch With Trainable Masked Layers
- ShrinkTeaNet: Million-scale Lightweight Face Recognition via Shrinking Teacher-Student Networks
- Bayesian Neural Network Priors Revisited
- Learned Threshold Pruning
- A Closer Look at Structured Pruning for Neural Network Compression
- Resource-Efficient Neural Networks for Embedded Systems
- Adversarial Neural Pruning with Latent Vulnerability Suppression
- A Brain-inspired Algorithm for Training Highly Sparse Neural Networks
- SeReNe: Sensitivity based Regularization of Neurons for Structured Sparsity in Neural Networks
- Do We Actually Need Dense Over-Parameterization? In-Time Over-Parameterization in Sparse Training
- -ARM: Network Sparsification via Stochastic Binary Optimization
- How Much Can I Trust You? -- Quantifying Uncertainties in Explaining Neural Networks
- Learning Global Pairwise Interactions with Bayesian Neural Networks
- Learning Sparse Sharing Architectures for Multiple Tasks
- Robust Sparse Regularization: Simultaneously Optimizing Neural Network Robustness and Compactness
- Attended Temperature Scaling: A Practical Approach for Calibrating Deep Neural Networks
- Sampling-Free Variational Inference of Bayesian Neural Networks by Variance Backpropagation
- On the Compression of Neural Networks Using -Norm Regularization and Weight Pruning
- Dynamic Hard Pruning of Neural Networks at the Edge of the Internet
- Taxonomy and Evaluation of Structured Compression of Convolutional Neural Networks
- Explaining Bayesian Neural Networks
- On Convergence and Generalization of Dropout Training
- One-Shot Pruning of Recurrent Neural Networks by Jacobian Spectrum Evaluation
- Observation Space Matters: Benchmark and Optimization Algorithm
- Sparse Training via Boosting Pruning Plasticity with Neuroregeneration
- LocalDrop: A Hybrid Regularization for Deep Neural Networks
- Neural network relief: a pruning algorithm based on neural activity
- GNN is a Counter? Revisiting GNN for Question Answering
- Directional Pruning of Deep Neural Networks
- Layer Adaptive Node Selection in Bayesian Neural Networks: Statistical Guarantees and Implementation Details
- Deep Evidential Regression
- Contextual Dropout: An Efficient Sample-Dependent Dropout Module
- AC/DC: Alternating Compressed/DeCompressed Training of Deep Neural Networks
- Not All Attention Is All You Need
- Hierarchical Indian Buffet Neural Networks for Bayesian Continual Learning
- Evaluating the Robustness of Bayesian Neural Networks Against Different Types of Attacks
- Sparse Weight Activation Training
- HEMP: High-order Entropy Minimization for neural network comPression
- Growing Efficient Deep Networks by Structured Continuous Sparsification
- Neural 3D Scene Compression via Model Compression
- Radial and Directional Posteriors for Bayesian Neural Networks
- SparseRT: Accelerating Unstructured Sparsity on GPUs for Deep Learning Inference
- DeepCABAC: Context-adaptive binary arithmetic coding for deep neural network compression
- Campfire: Compressible, Regularization-Free, Structured Sparse Training for Hardware Accelerators
- Enabling Retrain-free Deep Neural Network Pruning using Surrogate Lagrangian Relaxation
- Minimizing FLOPs to Learn Efficient Sparse Representations
- Powerpropagation: A sparsity inducing weight reparameterisation
- Calibrate and Prune: Improving Reliability of Lottery Tickets Through Prediction Calibration
- VINNAS: Variational Inference-based Neural Network Architecture Search
- Sparse evolutionary Deep Learning with over one million artificial neurons on commodity hardware
- Non-Differentiable Supervised Learning with Evolution Strategies and Hybrid Methods
- Revisiting Loss Modelling for Unstructured Pruning
- Weight Pruning via Adaptive Sparsity Loss
- Exploiting Channel Similarity for Accelerating Deep Convolutional Neural Networks
- Landscape Connectivity and Dropout Stability of SGD Solutions for Over-parameterized Neural Networks
- Distributed Weight Consolidation: A Brain Segmentation Case Study
- Pruning Randomly Initialized Neural Networks with Iterative Randomization
- Multi-Task Variational Information Bottleneck
- Network Pruning for Low-Rank Binary Indexing
- How Large a Vocabulary Does Text Classification Need? A Variational Approach to Vocabulary Selection
- Trends and Advancements in Deep Neural Network Communication
- Walsh-Hadamard Variational Inference for Bayesian Deep Learning
- Topological Insights into Sparse Neural Networks
- Connectivity Matters: Neural Network Pruning Through the Lens of Effective Sparsity
- Lottery Jackpots Exist in Pre-trained Models
- Multi-Objective Pruning for CNNs Using Genetic Algorithm
- Informative Bayesian Neural Network Priors for Weak Signals
- Effective Sparsification of Neural Networks with Global Sparsity Constraint
- Neural Pruning via Growing Regularization
- A Winning Hand: Compressing Deep Networks Can Improve Out-Of-Distribution Robustness
- Learning Sparse & Ternary Neural Networks with Entropy-Constrained Trained Ternarization (EC2T)
- Combining Model and Parameter Uncertainty in Bayesian Neural Networks
- Principal Component Networks: Parameter Reduction Early in Training
- Greedy Optimization Provably Wins the Lottery: Logarithmic Number of Winning Tickets is Enough
- Deep Neural Compression Via Concurrent Pruning and Self-Distillation
- Learning Task-Oriented Communication for Edge Inference: An Information Bottleneck Approach
- SparseDNN: Fast Sparse Deep Learning Inference on CPUs
- SASL: Saliency-Adaptive Sparsity Learning for Neural Network Acceleration
- Image Captioning with Sparse Recurrent Neural Network
- Predefined Sparseness in Recurrent Sequence Models
- Nonparametric Bayesian Deep Networks with Local Competition
- Structured Compression by Weight Encryption for Unstructured Pruning and Quantization
- Knowing what you know in brain segmentation using Bayesian deep neural networks
- Why Lottery Ticket Wins? A Theoretical Perspective of Sample Complexity on Pruned Neural Networks
- Structured Dropout Variational Inference for Bayesian Neural Networks
- Beyond Dropout: Feature Map Distortion to Regularize Deep Neural Networks
- Joint Channel and Weight Pruning for Model Acceleration on Moblie Devices
- Variational Resampling Based Assessment of Deep Neural Networks under Distribution Shift
- Training Sparse Neural Networks using Compressed Sensing
- Dirichlet Pruning for Neural Network Compression
- Variational Bayesian Dropout with a Hierarchical Prior
- Mining the Weights Knowledge for Optimizing Neural Network Structures
- Dynamic Narrowing of VAE Bottlenecks Using GECO and L0 Regularization
- The Role of Regularization in Shaping Weight and Node Pruning Dependency and Dynamics
- Adaptive Variational Bayesian Inference for Sparse Deep Neural Network
- Probabilistic Approach for Road-Users Detection
- Gaussian Mean Field Regularizes by Limiting Learned Information
- Spectral Pruning for Recurrent Neural Networks
- Group Pruning using a Bounded-Lp norm for Group Gating and Regularization
- The Flip Side of the Reweighted Coin: Duality of Adaptive Dropout and Regularization
- Efficient Approximate Inference with Walsh-Hadamard Variational Inference
- Sparse within Sparse Gaussian Processes using Neighbor Information
- A Note on Latency Variability of Deep Neural Networks for Mobile Inference
- NodeDrop: A Condition for Reducing Network Size without Effect on Output
- Bayesian Sparsification Methods for Deep Complex-valued Networks
- Extracting representations of cognition across neuroimaging studies improves brain decoding
- Explore the Knowledge contained in Network Weights to Obtain Sparse Neural Networks
- A Discriminative Gaussian Mixture Model with Sparsity
- Modulating Regularization Frequency for Efficient Compression-Aware Model Training
- Evidential Turing Processes
- Distilling Knowledge From a Deep Pose Regressor Network
- How Well Do Sparse Imagenet Models Transfer?
- Convergence of a Relaxed Variable Splitting Coarse Gradient Descent Method for Learning Sparse Weight Binarized Activation Neural Networks
- Learning Sparsity of Representations with Discrete Latent Variables
- CODA: Constructivism Learning for Instance-Dependent Dropout Architecture Construction
- Dep-: Improving -based Network Sparsification via Dependency Modeling
- Probabilistic fine-tuning of pruning masks and PAC-Bayes self-bounded learning
- Joint Regularization on Activations and Weights for Efficient Neural Network Pruning
- Non-Convex Compressed Sensing with Training Data
- Cascade Weight Shedding in Deep Neural Networks: Benefits and Pitfalls for Network Pruning
- Out-of-the-box channel pruned networks
- PipeTune: Pipeline Parallelism of Hyper and System Parameters Tuning for Deep Learning Clusters
- Channel selection using Gumbel Softmax
- A Novel Unsupervised Post-Processing Calibration Method for DNNS with Robustness to Domain Shift