Representation Learning for Online and Offline RL in Low-rank MDPs
arXiv:2110.04652
Abstract
This work studies the question of Representation Learning in RL: how can we learn a compact low-dimensional representation such that on top of the representation we can perform RL procedures such as exploration and exploitation, in a sample efficient manner. We focus on the low-rank Markov Decision Processes (MDPs) where the transition dynamics correspond to a low-rank transition matrix. Unlike prior works that assume the representation is known (e.g., linear MDPs), here we need to learn the representation for the low-rank MDP. We study both the online RL and offline RL settings. For the online setting, operating with the same computational oracles used in FLAMBE (Agarwal et.al), the state-of-art algorithm for learning representations in low-rank MDPs, we propose an algorithm REP-UCB Upper Confidence Bound driven Representation learning for RL), which significantly improves the sample complexity from for FLAMBE to with being the rank of the transition matrix (or dimension of the ground truth representation), being the number of actions, and being the discounted factor. Notably, REP-UCB is simpler than FLAMBE, as it directly balances the interplay between representation learning, exploration, and exploitation, while FLAMBE is an explore-then-commit style approach and has to perform reward-free exploration step-by-step forward in time. For the offline RL setting, we develop an algorithm that leverages pessimism to learn under a partial coverage condition: our algorithm is able to compete against any policy as long as it is covered by the offline distribution.
References in corpus (28)
- Conservative Q-Learning for Offline Reinforcement Learning
- MOPO: Model-based Offline Policy Optimization
- Contextual Decision Processes with Low Bellman Rank are PAC-Learnable
- Learning Near Optimal Policies with Low Inherent Bellman Error
- Information-Theoretic Considerations in Batch Reinforcement Learning
- What are the Statistical Limits of Offline RL with Linear Function Approximation?
- FLAMBE: Structural Complexity and Representation Learning of Low Rank MDPs
- Is a Good Representation Sufficient for Sample Efficient Reinforcement Learning?
- Active Learning for Nonlinear System Identification with Guarantees
- PC-PG: Policy Cover Directed Exploration for Provable Policy Gradient Learning
- Information Theoretic Regret Bounds for Online Nonlinear Control
- Is Pessimism Provably Efficient for Offline RL?
- Instance-Dependent Complexity of Contextual Bandits and Reinforcement Learning: A Disagreement-Based Perspective
- Representation Matters: Offline Pretraining for Sequential Decision Making
- Near-Optimal Offline Reinforcement Learning via Double Variance Reduction
- Bellman-consistent Pessimism for Offline Reinforcement Learning
- Provable Benefits of Actor-Critic Methods for Offline Reinforcement Learning
- Model-free Representation Learning and Exploration in Low-rank MDPs
- Efficient Model-Based Reinforcement Learning through Optimistic Policy Search and Planning
- Leveraging Good Representations in Linear Contextual Bandits
- Mitigating Covariate Shift in Imitation Learning via Offline Data Without Great Coverage
- Cautiously Optimistic Policy Optimization and Exploration with Linear Function Approximation
- Corruption-Robust Offline Reinforcement Learning
- On the Power of Multitask Representation Learning in Linear MDP
- PC-MLP: Model-based Reinforcement Learning with Policy Cover Guided Exploration
- Provably Efficient Representation Selection in Low-rank Markov Decision Processes: From Online to Offline RL
- Learning Good State and Action Representations via Tensor Decomposition
- Agnostic Reinforcement Learning with Low-Rank MDPs and Rich Observations