DeepMDP: Learning Continuous Latent Space Models for Representation Learning
arXiv:1906.02736
Abstract
Many reinforcement learning (RL) tasks provide the agent with high-dimensional observations that can be simplified into low-dimensional continuous states. To formalize this process, we introduce the concept of a DeepMDP, a parameterized latent space model that is trained via the minimization of two tractable losses: prediction of rewards and prediction of the distribution over next latent states. We show that the optimization of these objectives guarantees (1) the quality of the latent space as a representation of the state space and (2) the quality of the DeepMDP as a model of the environment. We connect these results to prior work in the bisimulation literature, and explore the use of a variety of metrics. Our theoretical findings are substantiated by the experimental result that a trained DeepMDP recovers the latent structure underlying high-dimensional observations on a synthetic environment. Finally, we show that learning a DeepMDP as an auxiliary task in the Atari 2600 domain leads to large performance improvements over model-free RL.
13 pages main text, 16 pages appendix. ICML 2019
References in corpus (7)
- Learning to Navigate in Complex Environments
- The Cramer Distance as a Solution to Biased Wasserstein Gradients
- Dopamine: A Research Framework for Deep Reinforcement Learning
- Distributional Reinforcement Learning with Quantile Regression
- Visual Interaction Networks
- Hyperbolic Discounting and Learning over Multiple Horizons
- A Comparative Analysis of Expected and Distributional Reinforcement Learning
Cited by in corpus (15)
- Contrastive Learning of Structured World Models
- A Geometric Perspective on Optimal Representations for Reinforcement Learning
- Discriminative Particle Filter Reinforcement Learning for Complex Partial Observations
- Return-Based Contrastive Representation Learning for Reinforcement Learning
- Offline Reinforcement Learning from Images with Latent Space Models
- The Value Equivalence Principle for Model-Based Reinforcement Learning
- Which Mutual-Information Representation Learning Objectives are Sufficient for Control?
- Revisit Recommender System in the Permutation Prospective
- GRN: Generative Rerank Network for Context-wise Recommendation
- Models, Pixels, and Rewards: Evaluating Design Trade-offs in Visual Model-Based Reinforcement Learning
- Towards Robust Bisimulation Metric Learning
- Low-Dimensional State and Action Representation Learning with MDP Homomorphism Metrics
- Low-Variance Policy Gradient Estimation with World Models
- A Geometric Perspective on Self-Supervised Policy Adaptation
- A Self-Supervised Auxiliary Loss for Deep RL in Partially Observable Settings