Invariant Causal Prediction for Block MDPs
arXiv:2003.06016
Abstract
Generalization across environments is critical to the successful application of reinforcement learning algorithms to real-world challenges. In this paper, we consider the problem of learning abstractions that generalize in block MDPs, families of environments with a shared latent state space and dynamics structure over that latent space, but varying observations. We leverage tools from causal inference to propose a method of invariant prediction to learn model-irrelevance state abstractions (MISA) that generalize to novel observations in the multi-environment setting. We prove that for certain classes of environments, this approach outputs with high probability a state abstraction corresponding to the causal feature set with respect to the return. We further provide more general bounds on model error and generalization error in the multi-environment setting, in the process showing a connection between causal variable selection and the state abstraction framework for MDPs. We give empirical evidence that our methods work in both linear and nonlinear settings, attaining improved generalization over single- and multi-task baselines.
Accepted to ICML 2020. 16 pages, 8 figures
References in corpus (11)
- DeepMind Control Suite
- Adversarial Discriminative Domain Adaptation
- Quantifying Generalization in Reinforcement Learning
- Distral: Robust Multitask Reinforcement Learning
- Spectrally-normalized margin bounds for neural networks
- Algorithmic Framework for Model-based Deep Reinforcement Learning with Theoretical Guarantees
- On Causal and Anticausal Learning
- Meta-Learning without Memorization
- DeepMDP: Learning Continuous Latent Space Models for Representation Learning
- Sharing Knowledge in Multi-Task Deep Reinforcement Learning
- Natural Environment Benchmarks for Reinforcement Learning
Cited by in corpus (18)
- A Survey of Zero-shot Generalisation in Deep Reinforcement Learning
- Learning Invariant Representations for Reinforcement Learning without Reconstruction
- Automatic Data Augmentation for Generalization in Deep Reinforcement Learning
- Multi-Agent Reinforcement Learning: Methods, Applications, Visionary Prospects, and Challenges
- Structure in Deep Reinforcement Learning: A Survey and Open Problems
- Out-of-distribution Prediction with Invariant Risk Minimization: The Limitation and An Effective Fix
- Learning Task-relevant Representations for Generalization via Characteristic Functions of Reward Sequence Distributions
- Invariant Policy Optimization: Towards Stronger Generalization in Reinforcement Learning
- Intervention Design for Effective Sim2Real Transfer
- Reimagining an autonomous vehicle
- Environment Invariant Linear Least Squares
- Decoupling Value and Policy for Generalization in Reinforcement Learning
- Model-Invariant State Abstractions for Model-Based Reinforcement Learning
- Optimization-based Causal Estimation from Heterogenous Environments
- Towards Robust Off-Policy Evaluation via Human Inputs
- Building Minimal and Reusable Causal State Abstractions for Reinforcement Learning
- Learning Representations for Pixel-based Control: What Matters and Why?
- Effect-Invariant Mechanisms for Policy Generalization