2 papers
cs.LG2025
Average-DICE: Stationary Distribution Correction by Regression
Fengdi Che, Bryan Chan, Chen Ma +1
Off-policy policy evaluation (OPE), an essential component of reinforcement learning, has long suffered from stationary state distribution mismatch, undermining both stability and…
cs.LG2024
Provably Efficient Exploration in Reward Machines with Low Regret
Hippolyte Bourel, Anders Jonsson, Odalric-Ambrym Maillard +2
We study reinforcement learning (RL) for decision processes with non-Markovian reward, in which high-level knowledge of the task in the form of reward machines is available to the…