Learning Invariant Representations for Reinforcement Learning without Reconstruction
arXiv:2006.10742
Abstract
We study how representation learning can accelerate reinforcement learning from rich observations, such as images, without relying either on domain knowledge or pixel-reconstruction. Our goal is to learn representations that both provide for effective downstream control and invariance to task-irrelevant details. Bisimulation metrics quantify behavioral similarity between states in continuous MDPs, which we propose using to learn robust latent representations which encode only the task-relevant information from observations. Our method trains encoders such that distances in latent space equal bisimulation distances in state space. We demonstrate the effectiveness of our method at disregarding task-irrelevant information using modified visual MuJoCo tasks, where the background is replaced with moving distractors and natural videos, while achieving SOTA performance. We also test a first-person highway driving task where our method learns invariance to clouds, weather, and time of day. Finally, we provide generalization results drawn from properties of bisimulation metrics, and links to causal inference.
Accepted as an oral at ICLR 2021
References in corpus (11)
- Representation Learning with Contrastive Predictive Coding
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
- The Kinetics Human Action Video Dataset
- Data-Efficient Image Recognition with Contrastive Predictive Coding
- DeepMind Control Suite
- Learning Latent Dynamics for Planning from Pixels
- Efficient Off-Policy Meta-Reinforcement Learning via Probabilistic Context Variables
- Causality for Machine Learning
- Provably efficient RL with Rich Observations via Latent State Decoding
- Invariant Causal Prediction for Block MDPs
- Natural Environment Benchmarks for Reinforcement Learning
Cited by in corpus (32)
- A Survey of Zero-shot Generalisation in Deep Reinforcement Learning
- Automatic Data Augmentation for Generalization in Deep Reinforcement Learning
- Stabilizing Deep Q-Learning with ConvNets and Vision Transformers under Data Augmentation
- Representation Matters: Offline Pretraining for Sequential Decision Making
- Return-Based Contrastive Representation Learning for Reinforcement Learning
- Structure in Deep Reinforcement Learning: A Survey and Open Problems
- SECANT: Self-Expert Cloning for Zero-Shot Generalization of Visual Policies
- The Distracting Control Suite -- A Challenging Benchmark for Reinforcement Learning from Pixels
- Model-free Representation Learning and Exploration in Low-rank MDPs
- High-Dimensional Bayesian Optimisation with Variational Autoencoders and Deep Metric Learning
- Why Generalization in RL is Difficult: Epistemic POMDPs and Implicit Partial Observability
- Dreaming: Model-based Reinforcement Learning by Latent Imagination without Reconstruction
- Measuring Visual Generalization in Continuous Control from Pixels
- HALMA: Humanlike Abstraction Learning Meets Affordance in Rapid Problem Solving
- Intervention Design for Effective Sim2Real Transfer
- Provable Representation Learning for Imitation with Contrastive Fourier Features
- Domain Adversarial Reinforcement Learning
- AdaRL: What, Where, and How to Adapt in Transfer Reinforcement Learning
- Model-Based Domain Generalization
- Jointly-Learned State-Action Embedding for Efficient Reinforcement Learning
- TRAIL: Near-Optimal Imitation Learning with Suboptimal Data
- Learning Robust State Abstractions for Hidden-Parameter Block MDPs
- Evaluating the progress of Deep Reinforcement Learning in the real world: aligning domain-agnostic and domain-specific research
- Function Contrastive Learning of Transferable Meta-Representations
- Agent Modelling under Partial Observability for Deep Reinforcement Learning
- Decoupling Value and Policy for Generalization in Reinforcement Learning
- Learning State Representations via Retracing in Reinforcement Learning
- Common Information based Approximate State Representations in Multi-Agent Reinforcement Learning
- Towards Deeper Deep Reinforcement Learning with Spectral Normalization
- Provable RL with Exogenous Distractors via Multistep Inverse Dynamics
- A Geometric Perspective on Self-Supervised Policy Adaptation
- Semi-supervised Learning for Dense Object Detection in Retail Scenes