Sticking the Landing: Simple, Lower-Variance Gradient Estimators for Variational Inference
arXiv:1703.09194
Abstract
We propose a simple and general variant of the standard reparameterized gradient estimator for the variational evidence lower bound. Specifically, we remove a part of the total derivative with respect to the variational parameters that corresponds to the score function. Removing this term produces an unbiased gradient estimator whose variance approaches zero as the approximate posterior approaches the exact posterior. We analyze the behavior of this gradient estimator theoretically and empirically, and generalize it to more complex variational distributions such as mixtures and importance-weighted posteriors.
Cited by in corpus (36)
- An Introduction to Variational Autoencoders
- NVAE: A Deep Hierarchical Variational Autoencoder
- Educating Text Autoencoders: Latent Representation Guidance via Denoising
- Gaussian variational approximation for high-dimensional state space models
- Estimating Gradients for Discrete Random Variables by Sampling without Replacement
- A Review of Learning with Deep Generative Models from Perspective of Graphical Modeling
- Sampling-Free Variational Inference of Bayesian Neural Networks by Variance Backpropagation
- Infinitely Deep Bayesian Neural Networks with Stochastic Differential Equations
- Posterior inference unchained with EL_2O
- Kalman Gradient Descent: Adaptive Variance Reduction in Stochastic Optimization
- Leveraging the Exact Likelihood of Deep Latent Variable Models
- Independent Subspace Analysis for Unsupervised Learning of Disentangled Representations
- Mutual Information Gradient Estimation for Representation Learning
- An Easy to Interpret Diagnostic for Approximate Inference: Symmetric Divergence Over Simulations
- Failure Modes of Variational Autoencoders and Their Effects on Downstream Tasks
- Augment and Reduce: Stochastic Inference for Large Categorical Distributions
- On importance-weighted autoencoders
- GO Gradient for Expectation-Based Objectives
- On the Difficulty of Unbiased Alpha Divergence Minimization
- Variational Inference with Holder Bounds
- Invertible Gaussian Reparameterization: Revisiting the Gumbel-Softmax
- Automatic Differentiation Variational Inference with Mixtures
- Monte Carlo Filtering Objectives: A New Family of Variational Objectives to Learn Generative Model and Neural Adaptive Proposal for Time Series
- New Tricks for Estimating Gradients of Expectations
- Variational Inference over Non-differentiable Cardiac Simulators using Bayesian Optimization
- Generalized Doubly Reparameterized Gradient Estimators
- CRAUM-Net: Contextual Recursive Attention with Uncertainty Modeling for Salient Object Detection
- Amortized variance reduction for doubly stochastic objectives
- Pathwise Derivatives for Multivariate Distributions
- The equivalence between Stein variational gradient descent and black-box variational inference
- Mutual Information Constraints for Monte-Carlo Objectives
- Quantized Variational Inference
- A variational approximate posterior for the deep Wishart process
- Challenging the Semi-Supervised VAE Framework for Text Classification
- Stochastic Bayesian Neural Networks
- Controlling the Interaction Between Generation and Inference in Semi-Supervised Variational Autoencoders Using Importance Weighting