3 papers
cs.LG2025
A Unifying View of Linear Function Approximation in Off-Policy RL Through Matrix Splitting and Preconditioning
Zechen Wu, Amy Greenwald, Ronald Parr
In off-policy policy evaluation (OPE) tasks within reinforcement learning, Temporal Difference Learning(TD) and Fitted Q-Iteration (FQI) have traditionally been viewed as differing…
cs.LG2024
Mitigating Partial Observability in Sequential Decision Processes via the Lambda Discrepancy
Cameron Allen, Aaron Kirtland, Ruo Yu Tao +7
Reinforcement learning algorithms typically rely on the assumption that the environment dynamics and value function can be expressed in terms of a Markovian state representation. H…
cs.LG2024
An Optimal Tightness Bound for the Simulation Lemma
Sam Lobel, Ronald Parr
We present a bound for value-prediction error with respect to model misspecification that is tight, including constant factors. This is a direct improvement of the "simulation lemm…