Variational inference for Monte Carlo objectives
arXiv:1602.06725
Abstract
Recent progress in deep latent variable models has largely been driven by the development of flexible and scalable variational inference methods. Variational training of this type involves maximizing a lower bound on the log-likelihood, using samples from the variational posterior to compute the required gradients. Recently, Burda et al. (2016) have derived a tighter lower bound using a multi-sample importance sampling estimate of the likelihood and showed that optimizing it yields models that use more of their capacity and achieve higher likelihoods. This development showed the importance of such multi-sample objectives and explained the success of several related approaches. We extend the multi-sample approach to discrete latent variables and analyze the difficulty encountered when estimating the gradients involved. We then develop the first unbiased gradient estimator designed for importance-sampled objectives and evaluate it at training generative and structured output prediction models. The resulting estimator, which is based on low-variance per-sample learning signals, is both simpler and more effective than the NVIL estimator proposed for the single-sample variational objective, and is competitive with the currently used biased estimators.
Appears in Proceedings of the 33rd International Conference on Machine Learning (ICML), New York, NY, USA, 2016. JMLR: W&CP volume 48
References in corpus (4)
Cited by in corpus (72)
- An Introduction to Variational Autoencoders
- Disentangling by Factorising
- Wireless Image Transmission Using Deep Source Channel Coding With Attention Modules
- Recent Advances in Autoencoder-Based Representation Learning
- Fast Decoding in Sequence Models using Discrete Latent Variables
- Boundary-Seeking Generative Adversarial Networks
- Unsupervised Learning of 3D Structure from Images
- Regularizing Deep Neural Networks by Noise: Its Interpretation and Optimization
- Discrete Sequential Prediction of Continuous Actions for Deep RL
- Discrete Variational Autoencoders
- DVAE++: Discrete Variational Autoencoders with Overlapping Transformations
- ZhuSuan: A Library for Bayesian Deep Learning
- A Tutorial on Deep Latent Variable Models of Natural Language
- Neural Nearest Neighbors Networks
- A Better Variant of Self-Critical Sequence Training
- Accurate and Diverse Sampling of Sequences based on a "Best of Many" Sample Objective
- Fast amortized inference of neural activity from calcium imaging data with variational autoencoders
- Variational Memory Addressing in Generative Models
- A Tandem Learning Rule for Effective Training and Rapid Inference of Deep Spiking Neural Networks
- Deep Gaussian Process-Based Bayesian Inference for Contaminant Source Localization
- A Review of Learning with Deep Generative Models from Perspective of Graphical Modeling
- Associative Compression Networks for Representation Learning
- Joint Distributions for TensorFlow Probability
- Adaptive Path-Integral Autoencoder: Representation Learning and Planning for Dynamical Systems
- Analysis of diversity-accuracy tradeoff in image captioning
- Efficient Marginalization of Discrete and Structured Latent Variables via Sparsity
- Optimizing Sequential Medical Treatments with Auto-Encoding Heuristic Search in POMDPs
- Reparameterization Gradient for Non-differentiable Models
- DisARM: An Antithetic Gradient Estimator for Binary Latent Variables
- Learning Discrete Structured Representations by Adversarially Maximizing Mutual Information
- Storchastic: A Framework for General Stochastic Automatic Differentiation
- Learning Interpretable Deep State Space Model for Probabilistic Time Series Forecasting
- Learning Deep Generative Models with Annealed Importance Sampling
- Amortized Bethe Free Energy Minimization for Learning MRFs
- Unsupervised Recurrent Neural Network Grammars
- Improved Gradient-Based Optimization Over Discrete Distributions
- Direct Optimization through for Discrete Variational Auto-Encoder
- Importance Weighted Adversarial Variational Autoencoders for Spike Inference from Calcium Imaging Data
- Select and Attend: Towards Controllable Content Selection in Text Generation
- The Thermodynamic Variational Objective
- Annealed Flow Transport Monte Carlo
- Probabilistic Mixture-of-Experts for Efficient Deep Reinforcement Learning
- Teaching deep neural networks to localize single molecules for super-resolution microscopy
- Probabilistic Circuits for Variational Inference in Discrete Graphical Models
- An online sequence-to-sequence model for noisy speech recognition
- Variational Inference with Holder Bounds
- VQ-DRAW: A Sequential Discrete VAE
- Stochastic Sequential Neural Networks with Structured Inference
- Infomax Neural Joint Source-Channel Coding via Adversarial Bit Flip
- Filtering Variational Objectives
- Neural Variational Inference and Learning in Undirected Graphical Models
- Optimal Variance Control of the Score Function Gradient Estimator for Importance Weighted Bounds
- Rao-Blackwellizing the Straight-Through Gumbel-Softmax Gradient Estimator
- Invertible Gaussian Reparameterization: Revisiting the Gumbel-Softmax
- Amortized Population Gibbs Samplers with Neural Sufficient Statistics
- On Signal-to-Noise Ratio Issues in Variational Inference for Deep Gaussian Processes
- Learning in Variational Autoencoders with Kullback-Leibler and Renyi Integral Bounds
- Semi-supervised Sequential Generative Models
- Approximate Inference in Discrete Distributions with Monte Carlo Tree Search and Value Functions
- Reliable Categorical Variational Inference with Mixture of Discrete Normalizing Flows
- Cooperative image captioning
- Neural Communication Systems with Bandwidth-limited Channel
- A Rule for Gradient Estimator Selection, with an Application to Variational Inference
- PAC-Bayes: Narrowing the Empirical Risk Gap in the Misspecified Bayesian Regime
- Double Control Variates for Gradient Estimation in Discrete Latent Variable Models
- Imitation Learning of Factored Multi-agent Reactive Models
- Lattice Representation Learning
- A Fourier View of REINFORCE
- Mutual Information Constraints for Monte-Carlo Objectives
- Entropy optimized semi-supervised decomposed vector-quantized variational autoencoder model based on transfer learning for multiclass text classification and generation
- Approximation Based Variance Reduction for Reparameterization Gradients
- Hybrid Memoised Wake-Sleep: Approximate Inference at the Discrete-Continuous Interface