1 paper
Sanath Kumar Krishnamurthy, Tanmay Gangwani, Sumeet Katariya +3
We consider the finite-horizon offline reinforcement learning (RL) setting, and are motivated by the challenge of learning the policy at any step h in dynamic programming (DP) algo…