FLAMBE: Structural Complexity and Representation Learning of Low Rank MDPs
arXiv:2006.10814
Abstract
In order to deal with the curse of dimensionality in reinforcement learning (RL), it is common practice to make parametric assumptions where values or policies are functions of some low dimensional feature space. This work focuses on the representation learning question: how can we learn such features? Under the assumption that the underlying (unknown) dynamics correspond to a low rank transition matrix, we show how the representation learning question is related to a particular non-linear matrix decomposition problem. Structurally, we make precise connections between these low rank MDPs and latent variable models, showing how they significantly generalize prior formulations for representation learning in RL. Algorithmically, we develop FLAMBE, which engages in exploration and representation learning for provably efficient RL in low rank transition models.
New algorithm and analysis to remove the reachability assumption
References in corpus (10)
- From -entropy to KL-entropy: Analysis of minimum information complexity density estimation
- Contextual Decision Processes with Low Bellman Rank are PAC-Learnable
- Provably Efficient Maximum Entropy Exploration
- Optimism in Reinforcement Learning with Generalized Linear Function Approximation
- Reinforcement Learning in Feature Space: Matrix Bandit, Kernels, and Regret Bound
- Comments on the Du-Kakade-Wang-Yang Lower Bounds
- Reward-Free Exploration for Reinforcement Learning
- Agnostic Q-learning with Function Approximation in Deterministic Systems: Tight Bounds on Approximation Error and Sample Complexity
- Learning with Good Feature Representations in Bandits and in RL with a Generative Model
- Contrastive estimation reveals topic posterior information to linear models
Cited by in corpus (5)
- Provable Benefits of Actor-Critic Methods for Offline Reinforcement Learning
- MADE: Exploration via Maximizing Deviation from Explored Regions
- The Power of Exploiter: Provable Multi-Agent RL in Large State Spaces
- Online Sparse Reinforcement Learning
- The Information Geometry of Unsupervised Reinforcement Learning