1 paper
Volodymyr Tkachuk, Gellért Weisz, Csaba Szepesvári
We consider offline reinforcement learning (RL) in H-horizon Markov decision processes (MDPs) under the linear qI¨-realizability assumption, where the action-value function of…