Gradient Estimation Using Stochastic Computation Graphs
arXiv:1506.05254
Abstract
In a variety of problems originating in supervised, unsupervised, and reinforcement learning, the loss function is defined by an expectation over a collection of random variables, which might be part of a probabilistic model or the external world. Estimating the gradient of this loss function, using samples, lies at the core of gradient-based learning algorithms for these problems. We introduce the formalism of stochastic computation graphs---directed acyclic graphs that include both deterministic functions and conditional probability distributions---and describe how to easily and automatically derive an unbiased estimator of the loss function's gradient. The resulting algorithm for computing the gradient estimator is a simple modification of the standard backpropagation algorithm. The generic scheme we propose unifies estimators derived in variety of prior work, along with variance-reduction techniques therein. It could assist researchers in developing intricate models involving a combination of stochastic and deterministic operations, enabling, for example, attention, memory, and control actions.
Advances in Neural Information Processing Systems 28 (NIPS 2015)
References in corpus (2)
Cited by in corpus (72)
- Continuous control with deep reinforcement learning
- A Brief Survey of Deep Reinforcement Learning
- TensorFlow Distributions
- Learning Discrete Structures for Graph Neural Networks
- Monte Carlo Gradient Estimation in Machine Learning
- Ordered Neurons: Integrating Tree Structures into Recurrent Neural Networks
- Anti-efficient encoding in emergent communication
- Pathwise Derivatives Beyond the Reparameterization Trick
- Learning Algorithms for Active Learning
- Seq2Slate: Re-ranking and Slate Optimization with RNNs
- Parallel Attention Network with Sequence Matching for Video Grounding
- Distilling Policy Distillation
- Stochastic Optimization of Sorting Networks via Continuous Relaxations
- MuProp: Unbiased Backpropagation for Stochastic Neural Networks
- VIREL: A Variational Inference Framework for Reinforcement Learning
- Learning Efficient Algorithms with Hierarchical Attentive Memory
- Automatic Differentiable Monte Carlo: Theory and Application
- DeepHealth: Review and challenges of artificial intelligence in health informatics
- Efficient Transformers with Dynamic Token Pooling
- ADEV: Sound Automatic Differentiation of Expected Values of Probabilistic Programs
- Estimating Gradients for Discrete Random Variables by Sampling without Replacement
- Learning Interpretable Deep Disentangled Neural Networks for Hyperspectral Unmixing
- Improving Exploration in Soft-Actor-Critic with Normalizing Flows Policies
- Reparameterization Gradients through Acceptance-Rejection Sampling Algorithms
- Learning Belief Representations for Imitation Learning in POMDPs
- Causal Modeling for Fairness in Dynamical Systems
- How to Learn a Useful Critic? Model-based Action-Gradient-Estimator Policy Optimization
- Functional Tensors for Probabilistic Programming
- Parallel Training of Deep Networks with Local Updates
- Strategic Maneuver and Disruption with Reinforcement Learning Approaches for Multi-Agent Coordination
- Storchastic: A Framework for General Stochastic Automatic Differentiation
- Probabilistic Programming with Programmable Variational Inference
- DynaNet: Neural Kalman Dynamical Model for Motion Estimation and Prediction
- Learning Proposals for Probabilistic Programs with Inference Combinators
- CommunityGAN: Community Detection with Generative Adversarial Nets
- Implicit MLE: Backpropagating Through Discrete Exponential Family Distributions
- Deep Adversarial Social Recommendation
- Deep clustering with concrete k-means
- Improved Gradient-Based Optimization Over Discrete Distributions
- Preventing Posterior Collapse with Levenshtein Variational Autoencoder
- Rescuing neural spike train models from bad MLE
- GO Gradient for Expectation-Based Objectives
- Loaded DiCE: Trading off Bias and Variance in Any-Order Score Function Estimators for Reinforcement Learning
- Entropy Minimization In Emergent Languages
- To Relieve Your Headache of Training an MRF, Take AdVIL
- A Stochastic Decoder for Neural Machine Translation
- Variance Reduction for Evolution Strategies via Structured Control Variates
- Backprop-Q: Generalized Backpropagation for Stochastic Computation Graphs
- Redistribution in Public Project Problems via Neural Networks
- Goal-directed Generation of Discrete Structures with Conditional Generative Models
- Rao-Blackwellizing the Straight-Through Gumbel-Softmax Gradient Estimator
- Semantics, Representations and Grammars for Deep Learning
- A unified view of likelihood ratio and reparameterization gradients and an optimal importance sampling scheme
- Zero-Shot Clinical Acronym Expansion via Latent Meaning Cells
- Cooperative image captioning
- Unifying Gradient Estimators for Meta-Reinforcement Learning via Off-Policy Evaluation
- Hierarchical Variational Imitation Learning of Control Programs
- Enforcing Reasoning in Visual Commonsense Reasoning
- Pathwise Derivatives for Multivariate Distributions
- Differentiable Antithetic Sampling for Variance Reduction in Stochastic Variational Inference
- No Representation without Transformation
- Generalized Transformation-based Gradient
- A Rule for Gradient Estimator Selection, with an Application to Variational Inference
- GO Hessian for Expectation-Based Objectives
- Factor Graph Grammars
- Learning a Deep Generative Model like a Program: the Free Category Prior
- Learning spatial hearing via innate mechanisms
- Lattice Representation Learning
- Accelerating Parameter Extraction of Power MOSFET Models Using Automatic Differentiation
- Perturbative estimation of stochastic gradients
- All-Action Policy Gradient Methods: A Numerical Integration Approach
- Hybrid Memoised Wake-Sleep: Approximate Inference at the Discrete-Continuous Interface