A Connection between Generative Adversarial Networks, Inverse Reinforcement Learning, and Energy-Based Models
arXiv:1611.03852
Abstract
Generative adversarial networks (GANs) are a recently proposed class of generative models in which a generator is trained to optimize a cost function that is being simultaneously learned by a discriminator. While the idea of learning cost functions is relatively new to the field of generative modeling, learning costs has long been studied in control and reinforcement learning (RL) domains, typically for imitation learning from demonstrations. In these fields, learning cost function underlying observed behavior is known as inverse reinforcement learning (IRL) or inverse optimal control. While at first the connection between cost learning in RL and cost learning in generative modeling may appear to be a superficial one, we show in this paper that certain IRL methods are in fact mathematically equivalent to GANs. In particular, we demonstrate an equivalence between a sample-based algorithm for maximum entropy IRL and a GAN in which the generator's density can be evaluated and is provided as an additional input to the discriminator. Interestingly, maximum entropy IRL is a special case of an energy-based model. We discuss the interpretation of GANs as an algorithm for training energy-based models, and relate this interpretation to other recent work that seeks to connect GANs and EBMs. By formally highlighting the connection between GANs, IRL, and EBMs, we hope that researchers in all three communities can better identify and apply transferable ideas from one domain to another, particularly for developing more stable and scalable algorithms: a major challenge in all three domains.
NIPS 2016 Workshop on Adversarial Training. First two authors contributed equally
References in corpus (5)
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- Energy-based Generative Adversarial Network
- An Actor-Critic Algorithm for Sequence Prediction
- Connecting Generative Adversarial Networks and Actor-Critic Methods
- Reward Augmented Maximum Likelihood for Neural Structured Prediction
Cited by in corpus (52)
- A Review on Generative Adversarial Networks: Algorithms, Theory, and Applications
- Deep Learning and Knowledge-Based Methods for Computer Aided Molecular Design -- Toward a Unified Approach: State-of-the-Art and Future Directions
- Connecting Generative Adversarial Networks and Actor-Critic Methods
- A deep inverse reinforcement learning approach to route choice modeling with context-dependent rewards
- Imitating Interactive Intelligence
- Offline Reinforcement Learning with Fisher Divergence Critic Regularization
- A Divergence Minimization Perspective on Imitation Learning Methods
- Learning Loss Functions for Semi-supervised Learning via Discriminative Adversarial Networks
- SafeCritic: Collision-Aware Trajectory Prediction
- Learning Energy-Based Models by Diffusion Recovery Likelihood
- Deep Bayesian Reward Learning from Preferences
- -GAIL: Learning -Divergence for Generative Adversarial Imitation Learning
- State-only Imitation with Transition Dynamics Mismatch
- Demonstration-Guided Reinforcement Learning with Learned Skills
- Weakly-supervised Knowledge Graph Alignment with Adversarial Learning
- Discriminator Contrastive Divergence: Semi-Amortized Generative Modeling by Exploring Energy of the Discriminator
- Energy-Based Sequence GANs for Recommendation and Their Connection to Imitation Learning
- Situated GAIL: Multitask imitation using task-conditioned adversarial inverse reinforcement learning
- Non-Adversarial Imitation Learning and its Connections to Adversarial Methods
- Learning Equivariant Energy Based Models with Equivariant Stein Variational Gradient Descent
- Trajectory Prediction with Latent Belief Energy-Based Model
- Identifiability in inverse reinforcement learning
- Off-Policy Adversarial Inverse Reinforcement Learning
- Domain-Robust Visual Imitation Learning with Mutual Information Constraints
- Robust Maximum Entropy Behavior Cloning
- f-IRL: Inverse Reinforcement Learning via State Marginal Matching
- Learning Multi-Task Transferable Rewards via Variational Inverse Reinforcement Learning
- Adversarial recovery of agent rewards from latent spaces of the limit order book
- Inverse reinforcement learning for autonomous navigation via differentiable semantic mapping and planning
- Enhanced Experience Replay Generation for Efficient Reinforcement Learning
- Environment Reconstruction with Hidden Confounders for Reinforcement Learning based Recommendation
- Discriminator Soft Actor Critic without Extrinsic Rewards
- Deep Consensus Learning
- A Relation Analysis of Markov Decision Process Frameworks
- Data Driven Aircraft Trajectory Prediction with Deep Imitation Learning
- Imitation Learning for Fashion Style Based on Hierarchical Multimodal Representation
- Reward function shape exploration in adversarial imitation learning: an empirical study
- Learning Variable Impedance Control via Inverse Reinforcement Learning for Force-Related Tasks
- Reparameterized Variational Divergence Minimization for Stable Imitation
- oIRL: Robust Adversarial Inverse Reinforcement Learning with Temporally Extended Actions
- Support-weighted Adversarial Imitation Learning
- Sample Efficient Social Navigation Using Inverse Reinforcement Learning
- Plan-Space State Embeddings for Improved Reinforcement Learning
- Off-Dynamics Inverse Reinforcement Learning from Hetero-Domain
- Multi-Agent Inverse Reinforcement Learning: Suboptimal Demonstrations and Alternative Solution Concepts
- An Inattention Model for Traveler Behavior with e-Coupons
- Pre-Training Transformers as Energy-Based Cloze Models
- Sequential Anomaly Detection using Inverse Reinforcement Learning
- CAMEO: Curiosity Augmented Metropolis for Exploratory Optimal Policies
- Scenario Generalization of Data-driven Imitation Models in Crowd Simulation
- Generalized Inverse Planning: Learning Lifted non-Markovian Utility for Generalizable Task Representation
- Generative adversarial training of product of policies for robust and adaptive movement primitives