2 papers
cs.LG2024
Mitigating Partial Observability in Sequential Decision Processes via the Lambda Discrepancy
Cameron Allen, Aaron Kirtland, Ruo Yu Tao +7
Reinforcement learning algorithms typically rely on the assumption that the environment dynamics and value function can be expressed in terms of a Markovian state representation. H…
cs.LG2024
An Optimal Tightness Bound for the Simulation Lemma
Sam Lobel, Ronald Parr
We present a bound for value-prediction error with respect to model misspecification that is tight, including constant factors. This is a direct improvement of the "simulation lemm…