A Variant of the Wang-Foster-Kakade Lower Bound for the Discounted Setting
arXiv:2011.01075
Abstract
Recently, Wang et al. (2020) showed a highly intriguing hardness result for batch reinforcement learning (RL) with linearly realizable value function and good feature coverage in the finite-horizon case. In this note we show that once adapted to the discounted setting, the construction can be simplified to a 2-state MDP with 1-dimensional features, such that learning is impossible even with an infinite amount of data.
Cited by in corpus (5)
- Offline RL Without Off-Policy Evaluation
- Finite Sample Analysis of Minimax Offline Reinforcement Learning: Completeness, Fast Rates and First-Order Efficiency
- Instabilities of Offline RL with Pre-Trained Neural Representation
- Infinite-Horizon Offline Reinforcement Learning with Linear Function Approximation: Curse of Dimensionality and Algorithm
- Offline Reinforcement Learning: Fundamental Barriers for Value Function Approximation