2 papers
cs.LG2026
Trajectory Data Suffices for Statistically Efficient Policy Evaluation in Fixed-Horizon Offline RL with Linear -Realizability and Concentrability
Volodymyr Tkachuk, Csaba Szepesvári, Xiaoqi Tan
We study finite-horizon offline reinforcement learning (RL) with function approximation for both policy evaluation and policy optimization. Prior work established that statisticall…
cs.LG2024
Trajectory Data Suffices for Statistically Efficient Learning in Offline RL with Linear -Realizability and Concentrability
Volodymyr Tkachuk, Gellért Weisz, Csaba Szepesvári
We consider offline reinforcement learning (RL) in -horizon Markov decision processes (MDPs) under the linear -realizability assumption, where the action-value function of…