Variational Dropout and the Local Reparameterization Trick
arXiv:1506.02557
Abstract
We investigate a local reparameterizaton technique for greatly reducing the variance of stochastic gradients for variational Bayesian inference (SGVB) of a posterior over model parameters, while retaining parallelizability. This local reparameterization translates uncertainty about global parameters into local noise that is independent across datapoints in the minibatch. Such parameterizations can be trivially parallelized and have variance that is inversely proportional to the minibatch size, generally leading to much faster convergence. Additionally, we explore a connection with dropout: Gaussian dropout objectives correspond to SGVB with local reparameterization, a scale-invariant prior and proportionally fixed posterior variance. Our method allows inference of more flexibly parameterized posteriors; specifically, we propose variational dropout, a generalization of Gaussian dropout where the dropout rates are learned, often leading to better models. The method is demonstrated through several experiments.
References in corpus (6)
- Improving neural networks by preventing co-adaptation of feature detectors
- Weight Uncertainty in Neural Networks
- Probabilistic Backpropagation for Scalable Learning of Bayesian Neural Networks
- Bayesian Posterior Sampling via Stochastic Gradient Fisher Scoring
- A Bayesian encourages dropout
- Fast Adaptive Weight Noise
Cited by in corpus (208)
- An Introduction to Variational Autoencoders
- The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
- Physics-Constrained Deep Learning for High-dimensional Surrogate Modeling and Uncertainty Quantification without Labeled Data
- Can You Trust Your Model's Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift
- Towards Robust Evaluations of Continual Learning
- A Simple Baseline for Bayesian Uncertainty in Deep Learning
- A Comprehensive guide to Bayesian Convolutional Neural Network with Variational Inference
- Causal Confusion in Imitation Learning
- Preconditioned Stochastic Gradient Langevin Dynamics for Deep Neural Networks
- Beyond Finite Layer Neural Networks: Bridging Deep Architectures and Numerical Differential Equations
- Evaluating Bayesian Deep Learning Methods for Semantic Segmentation
- Bayesian Recurrent Neural Networks
- Quantifying model form uncertainty in Reynolds-averaged turbulence models with Bayesian deep neural networks
- Practical Deep Learning with Bayesian Principles
- SDNet: Contextualized Attention-based Deep Network for Conversational Question Answering
- Variational Federated Multi-Task Learning
- DeepCABAC: A Universal Compression Algorithm for Deep Neural Networks
- Mixout: Effective Regularization to Finetune Large-scale Pretrained Language Models
- Structured and Efficient Variational Deep Learning with Matrix Gaussian Posteriors
- SpArSe: Sparse Architecture Search for CNNs on Resource-Constrained Microcontrollers
- A Survey on Epistemic (Model) Uncertainty in Supervised Learning: Recent Advances and Applications
- Pitfalls of In-Domain Uncertainty Estimation and Ensembling in Deep Learning
- Learning Sparse Networks Using Targeted Dropout
- Sequential Neural Likelihood: Fast Likelihood-free Inference with Autoregressive Flows
- Bayesian Batch Active Learning as Sparse Subset Approximation
- Advanced Dropout: A Model-free Methodology for Bayesian Dropout Optimization
- Critical Learning Periods in Deep Neural Networks
- Uncertainty Estimations by Softplus normalization in Bayesian Convolutional Neural Networks with Variational Inference
- EDropout: Energy-Based Dropout and Pruning of Deep Neural Networks
- Resource-efficient Deep Neural Networks for Automotive Radar Interference Mitigation
- Effective and Efficient Dropout for Deep Convolutional Neural Networks
- Restricting the Flow: Information Bottlenecks for Attribution
- Estimating Uncertainty in Neural Networks for Cardiac MRI Segmentation: A Benchmark Study
- A Unifying Bayesian View of Continual Learning
- Tensorized Embedding Layers for Efficient Model Compression
- Survey of Dropout Methods for Deep Neural Networks
- Reinforcement Learning Based Graph-to-Sequence Model for Natural Question Generation
- TapType: Ten-finger text entry on everyday surfaces via Bayesian inference
- Where is the Information in a Deep Neural Network?
- Compression of Recurrent Neural Networks for Efficient Language Modeling
- Optimizing Millions of Hyperparameters by Implicit Differentiation
- On the Expressiveness of Approximate Inference in Bayesian Neural Networks
- Confidence Calibration for Convolutional Neural Networks Using Structured Dropout
- Explicit Regularisation in Gaussian Noise Injections
- Calibration of Deep Probabilistic Models with Decoupled Bayesian Neural Networks
- How Good is the Bayes Posterior in Deep Neural Networks Really?
- Efficient and Scalable Bayesian Neural Nets with Rank-1 Factors
- Information Aware Max-Norm Dirichlet Networks for Predictive Uncertainty Estimation
- 'In-Between' Uncertainty in Bayesian Neural Networks
- Uncertainty as a Form of Transparency: Measuring, Communicating, and Using Uncertainty
- Can Unstructured Pruning Reduce the Depth in Deep Neural Networks?
- Uncertainty Quantification in Deep Learning for Safer Neuroimage Enhancement
- Decentralized Bayesian Learning over Graphs
- Instance-aware, Context-focused, and Memory-efficient Weakly Supervised Object Detection
- Bayesian Neural Network Priors Revisited
- Probabilistic Deep Learning to Quantify Uncertainty in Air Quality Forecasting
- Bayesian Neural Networks With Maximum Mean Discrepancy Regularization
- The k-tied Normal Distribution: A Compact Parameterization of Gaussian Mean Field Posteriors in Bayesian Neural Networks
- A Survey on Principles, Models and Methods for Learning from Irregularly Sampled Time Series
- Learnable Bernoulli Dropout for Bayesian Deep Learning
- Resource-Efficient Neural Networks for Embedded Systems
- Adversarial Neural Pruning with Latent Vulnerability Suppression
- Uncertainty Weighted Actor-Critic for Offline Reinforcement Learning
- Semi-Supervised Deep Learning for Multi-Tissue Segmentation from Multi-Contrast MRI
- Subspace Inference for Bayesian Deep Learning
- SeReNe: Sensitivity based Regularization of Neurons for Structured Sparsity in Neural Networks
- -ARM: Network Sparsification via Stochastic Binary Optimization
- How Much Can I Trust You? -- Quantifying Uncertainties in Explaining Neural Networks
- Learning Global Pairwise Interactions with Bayesian Neural Networks
- Bayesian Graph Neural Networks with Adaptive Connection Sampling
- Dynamic Hard Pruning of Neural Networks at the Edge of the Internet
- Bayesian Inference for Large Scale Image Classification
- Sampling-Free Variational Inference of Bayesian Neural Networks by Variance Backpropagation
- Taxonomy and Evaluation of Structured Compression of Convolutional Neural Networks
- Explaining Bayesian Neural Networks
- Global inducing point variational posteriors for Bayesian neural networks and deep Gaussian processes
- Supervised Uncertainty Quantification for Segmentation with Multiple Annotations
- Robustly representing uncertainty in deep neural networks through sampling
- Normalized Flat Minima: Exploring Scale Invariant Definition of Flat Minima for Neural Networks using PAC-Bayesian Analysis
- Randomized Value Functions via Multiplicative Normalizing Flows
- Variance Networks: When Expectation Does Not Meet Your Expectations
- Natural Question Generation with Reinforcement Learning Based Graph-to-Sequence Model
- TraDE: Transformers for Density Estimation
- The Hidden Uncertainty in a Neural Networks Activations
- GNN is a Counter? Revisiting GNN for Question Answering
- Benchmarking the Neural Linear Model for Regression
- Contextual Dropout: An Efficient Sample-Dependent Dropout Module
- Calibration of Model Uncertainty for Dropout Variational Inference
- Deep Evidential Regression
- Exploring Bayesian Deep Learning for Urgent Instructor Intervention Need in MOOC Forums
- Improving predictions of Bayesian neural nets via local linearization
- Hierarchical Indian Buffet Neural Networks for Bayesian Continual Learning
- Interval Neural Networks: Uncertainty Scores
- Full deep neural network training on a pruned weight budget
- Interpreting Deep Neural Networks Through Variable Importance
- Bayesian Graph Neural Networks for Molecular Property Prediction
- Do End-to-End Speech Recognition Models Care About Context?
- Influence of uncertainty estimation techniques on false-positive reduction in liver lesion detection
- Neural Density Estimation and Likelihood-free Inference
- Tied Hidden Factors in Neural Networks for End-to-End Speaker Recognition
- Radial and Directional Posteriors for Bayesian Neural Networks
- Using Large Ensembles of Control Variates for Variational Inference
- DeepCABAC: Context-adaptive binary arithmetic coding for deep neural network compression
- Bayesian Few-Shot Classification with One-vs-Each Pólya-Gamma Augmented Gaussian Processes
- Robust Deep Gaussian Processes
- Exploring Gradient Flow Based Saliency for DNN Model Compression
- LiBRe: A Practical Bayesian Approach to Adversarial Detection
- Efficient non-conjugate Gaussian process factor models for spike count data using polynomial approximations
- Bayesian Inference Forgetting
- VINNAS: Variational Inference-based Neural Network Architecture Search
- Wide Neural Networks with Bottlenecks are Deep Gaussian Processes
- Attribution Preservation in Network Compression for Reliable Network Interpretation
- Multi Layer Neural Networks as Replacement for Pooling Operations
- Stochastic Neural Network with Kronecker Flow
- LeMoNADe: Learned Motif and Neuronal Assembly Detection in calcium imaging videos
- Ranking over Regression for Bayesian Optimization and Molecule Selection
- UFO-BLO: Unbiased First-Order Bilevel Optimization
- Distributed Weight Consolidation: A Brain Segmentation Case Study
- Ensemble Model Patching: A Parameter-Efficient Variational Bayesian Neural Network
- Multi-Task Variational Information Bottleneck
- Multi-Stage Transfer Learning with an Application to Selection Process
- A Bit More Bayesian: Domain-Invariant Learning with Uncertainty
- Tractable Approximate Gaussian Inference for Bayesian Neural Networks
- Hierarchical Gaussian Process Priors for Bayesian Neural Network Weights
- Text Generation with Exemplar-based Adaptive Decoding
- How Large a Vocabulary Does Text Classification Need? A Variational Approach to Vocabulary Selection
- FastFusionNet: New State-of-the-Art for DAWNBench SQuAD
- Causal Mediation Analysis Leveraging Multiple Types of Summary Statistics Data
- M-FAC: Efficient Matrix-Free Approximations of Second-Order Information
- Posterior Meta-Replay for Continual Learning
- Gradient-EM Bayesian Meta-learning
- On Calibration of Mixup Training for Deep Neural Networks
- Informative Bayesian Neural Network Priors for Weak Signals
- Meta Dropout: Learning to Perturb Features for Generalization
- Walsh-Hadamard Variational Inference for Bayesian Deep Learning
- Effect of latent space distribution on the segmentation of images with multiple annotations
- Learning Task-Oriented Communication for Edge Inference: An Information Bottleneck Approach
- Prediction intervals for Deep Neural Networks
- Quantifying Model Uncertainty in Inverse Problems via Bayesian Deep Gradient Descent
- Learning Partially Known Stochastic Dynamics with Empirical PAC Bayes
- Image Captioning with Sparse Recurrent Neural Network
- Uncertainty-Aware Model Adaptation for Unsupervised Cross-Domain Object Detection
- IPOD: An Industrial and Professional Occupations Dataset and its Applications to Occupational Data Mining and Analysis
- Adaptive Bayesian Linear Regression for Automated Machine Learning
- Bayesian Neural Networks for Virtual Flow Metering: An Empirical Study
- Hidden Markov Neural Networks
- Asymmetric Heavy Tails and Implicit Bias in Gaussian Noise Injections
- Structured Dropout Variational Inference for Bayesian Neural Networks
- Knowing what you know in brain segmentation using Bayesian deep neural networks
- Stochastic Model Pruning via Weight Dropping Away and Back
- Beyond the Mean-Field: Structured Deep Gaussian Processes Improve the Predictive Uncertainties
- Smoothed Inference for Adversarially-Trained Models
- Variational Depth Search in ResNets
- Sample Efficient Reinforcement Learning via Model-Ensemble Exploration and Exploitation
- Estimating Predictive Uncertainty Under Program Data Distribution Shift
- Flying Through a Narrow Gap Using End-to-end Deep Reinforcement Learning Augmented with Curriculum Learning and Sim2Real
- Unifying Variational Inference and PAC-Bayes for Supervised Learning that Scales
- Calibrated Top-1 Uncertainty estimates for classification by score based models
- Colored Noise Injection for Training Adversarially Robust Neural Networks
- Stabilising priors for robust Bayesian deep learning
- Entropy-SGD optimizes the prior of a PAC-Bayes bound: Generalization properties of Entropy-SGD and data-dependent priors
- Inferring Evidence from Nested Sampling Data via Information Field Theory
- Real-Time Uncertainty Estimation in Computer Vision via Uncertainty-Aware Distribution Distillation
- Gaussian Mean Field Regularizes by Limiting Learned Information
- Deep Reinforcement Learning with Weighted Q-Learning
- Probability Paths and the Structure of Predictions over Time
- Variational Bayesian Dropout with a Hierarchical Prior
- Adversarial Dropout for Recurrent Neural Networks
- Relational Reasoning Network (RRN) for Anatomical Landmarking
- Dissecting Non-Vacuous Generalization Bounds based on the Mean-Field Approximation
- TyXe: Pyro-based Bayesian neural nets for Pytorch
- On regularization of gradient descent, layer imbalance and flat minima
- The Flip Side of the Reweighted Coin: Duality of Adaptive Dropout and Regularization
- Deep Neural Networks as Point Estimates for Deep Gaussian Processes
- Probabilistic Approach for Road-Users Detection
- M-estimation with the Trimmed l1 Penalty
- Bayesian Sparsification Methods for Deep Complex-valued Networks
- Deep Model Compression Via Two-Stage Deep Reinforcement Learning
- Stochastic Bayesian Neural Networks
- Adaptive Weight Decay for Deep Neural Networks
- Generative Particle Variational Inference via Estimation of Functional Gradients
- Evidential Turing Processes
- Facetron: A Multi-speaker Face-to-Speech Model based on Cross-modal Latent Representations
- Dropout Drops Double Descent
- Multi-Task Neural Processes
- Why Calibration Error is Wrong Given Model Uncertainty: Using Posterior Predictive Checks with Deep Learning
- Robustness Against Outliers For Deep Neural Networks By Gradient Conjugate Priors
- Debiasing a First-order Heuristic for Approximate Bi-level Optimization
- CODA: Constructivism Learning for Instance-Dependent Dropout Architecture Construction
- Residual Overfit Method of Exploration
- Improving Adversarial Robustness for Free with Snapshot Ensemble
- Statistical Guarantees for Transformation Based Models with Applications to Implicit Variational Inference
- Quantized Variational Inference
- Variational Bayes Neural Network: Posterior Consistency, Classification Accuracy and Computational Challenges
- Extracting representations of cognition across neuroimaging studies improves brain decoding
- A Bayesian Neural Network based on Dropout Regulation
- Sequential Anomaly Detection using Inverse Reinforcement Learning
- Bayesian Generative Models for Knowledge Transfer in MRI Semantic Segmentation Problems
- Deep Direct Likelihood Knockoffs
- Quantal synaptic dilution enhances sparse encoding and dropout regularisation in deep networks
- SIM: A Slot-Independent Neural Model for Dialogue State Tracking
- Improving Predictive Uncertainty Estimation using Dropout -- Hamiltonian Monte Carlo
- Reconsidering Analytical Variational Bounds for Output Layers of Deep Networks
- Regularising Deep Networks with Deep Generative Models
- Efficient Approximate Inference with Walsh-Hadamard Variational Inference
- A Divergence Bound for Hybrids of MCMC and Variational Inference and an Application to Langevin Dynamics and SGVI
- The Variational InfoMax Learning Objective
- Nearest-Neighbor Neural Networks for Geostatistics