Flipout: Efficient Pseudo-Independent Weight Perturbations on Mini-Batches
arXiv:1803.04386
Abstract
Stochastic neural net weights are used in a variety of contexts, including regularization, Bayesian neural nets, exploration in reinforcement learning, and evolution strategies. Unfortunately, due to the large number of weights, all the examples in a mini-batch typically share the same weight perturbation, thereby limiting the variance reduction effect of large mini-batches. We introduce flipout, an efficient method for decorrelating the gradients within a mini-batch by implicitly sampling pseudo-independent weight perturbations for each example. Empirically, flipout achieves the ideal linear variance reduction for fully connected networks, convolutional networks, and RNNs. We find significant speedups in training neural networks with multiplicative Gaussian perturbations. We show that flipout is effective at regularizing LSTMs, and outperforms previous methods. Flipout also enables us to vectorize evolution strategies: in our experiments, a single GPU with flipout can handle the same throughput as at least 40 CPU cores using existing methods, equivalent to a factor-of-4 cost reduction on Amazon Web Services.
Published as a conference paper at ICLR 2018
References in corpus (10)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Auto-Encoding Variational Bayes
- Recurrent Neural Network Regularization
- On the Convergence of Adam and Beyond
- Evolution Strategies as a Scalable Alternative to Reinforcement Learning
- Noisy Networks for Exploration
- Regularizing and Optimizing LSTM Language Models
- Parameter Space Noise for Exploration
- Bayesian Compression for Deep Learning
- Zoneout: Regularizing RNNs by Randomly Preserving Hidden Activations
Cited by in corpus (60)
- A Review of Uncertainty Quantification in Deep Learning: Techniques, Applications and Challenges
- Can You Trust Your Model's Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift
- Hands-on Bayesian Neural Networks -- a Tutorial for Deep Learning Users
- Deep Ensembles: A Loss Landscape Perspective
- Adversarial Weight Perturbation Helps Robust Generalization
- Variational Federated Multi-Task Learning
- Deeply Uncertain: Comparing Methods of Uncertainty Quantification in Deep Learning Algorithms
- Bayesian Layers: A Module for Neural Network Uncertainty
- Probabilistic Deep Learning for Real-Time Large Deformation Simulations
- Learning Sparse Networks Using Targeted Dropout
- A Systematic Comparison of Bayesian Deep Learning Robustness in Diabetic Retinopathy Tasks
- Parameters Estimation for the Cosmic Microwave Background with Bayesian Neural Networks
- Detection of Gravitational Waves Using Bayesian Neural Networks
- Per-Object Systematics using Deep-Learned Calibration
- Dynamical mass inference of galaxy clusters with neural flows
- Hyperparameter Ensembles for Robustness and Uncertainty Quantification
- Bayesian Deep Ensembles via the Neural Tangent Kernel
- Laplace Redux -- Effortless Bayesian Deep Learning
- Uncertainty-Aware Reliable Text Classification
- Flattening Sharpness for Dynamic Gradient Projection Memory Benefits Continual Learning
- The Relevance of Bayesian Layer Positioning to Model Uncertainty in Deep Bayesian Active Learning
- Action and Perception as Divergence Minimization
- The k-tied Normal Distribution: A Compact Parameterization of Gaussian Mean Field Posteriors in Bayesian Neural Networks
- Transferable Calibration with Lower Bias and Variance in Domain Adaptation
- Uncertainty-Aware Self-Supervised Learning of Spatial Perception Tasks
- Reliable training and estimation of variance networks
- An Empirical Study of Large-Batch Stochastic Gradient Descent with Structured Covariance Noise
- Evaluating the Robustness of Bayesian Neural Networks Against Different Types of Attacks
- Constraining cosmological parameters from N-body simulations with Variational Bayesian Neural Networks
- Delta-STN: Efficient Bilevel Optimization for Neural Networks using Structured Response Jacobians
- Uncertainty aware audiovisual activity recognition using deep Bayesian variational inference
- Bayesian Neural Networks at Scale: A Performance Analysis and Pruning Study
- Deep Probabilistic Models to Detect Data Poisoning Attacks
- Variational Autoencoder Kernel Interpretation and Selection for Classification
- BAR: Bayesian Activity Recognition using variational inference
- Mean-Field Approximation to Gaussian-Softmax Integral with Application to Uncertainty Estimation
- Measuring the Substructure Mass Power Spectrum of 23 SLACS Strong Galaxy-Galaxy Lenses with Convolutional Neural Networks
- Probabilistically-autoencoded horseshoe-disentangled multidomain item-response theory models
- Probabilistic Multi-Layer Perceptrons for Wind Farm Condition Monitoring
- A Modular Deep Learning Pipeline for Galaxy-Scale Strong Gravitational Lens Detection and Modeling
- Empirical Frequentist Coverage of Deep Learning Uncertainty Quantification Procedures
- Estimation of redshift and associated uncertainty of Fermi/LAT extra-galactic sources with Deep Learning
- Improving Uncertainty-Error Correspondence in Deep Bayesian Medical Image Segmentation
- Being a Bit Frequentist Improves Bayesian Neural Networks
- Robustness via Cross-Domain Ensembles
- Reliable Uncertainties for Bayesian Neural Networks using Alpha-divergences
- Safety and Robustness in Decision Making: Deep Bayesian Recurrent Neural Networks for Somatic Variant Calling in Cancer
- Closed Form Variational Objectives For Bayesian Neural Networks with a Single Hidden Layer
- Dissecting Non-Vacuous Generalization Bounds based on the Mean-Field Approximation
- Deep inference of simulated strong lenses in ground-based surveys
- A Survey on Assessing the Generalization Envelope of Deep Neural Networks: Predictive Uncertainty, Out-of-distribution and Adversarial Samples
- Uncertainty-Aware Rank-One MIMO Q Network Framework for Accelerated Offline Reinforcement Learning
- Potential and limitations of machine-learning approaches to inclusive determinations
- Scene Uncertainty and the Wellington Posterior of Deterministic Image Classifiers
- Pathologies in priors and inference for Bayesian transformers
- Defense Through Diverse Directions
- Deep Bayesian Recurrent Neural Networks for Somatic Variant Calling in Cancer
- Loss convergence in a causal Bayesian neural network of retail firm performance
- A Novel Unsupervised Post-Processing Calibration Method for DNNS with Robustness to Domain Shift
- Selective Probabilistic Classifier Based on Hypothesis Testing