1 paper
Teodor V. Marinov, Alekh Agarwal, Mircea Trofin
This work studies a Reinforcement Learning (RL) problem in which we are given a set of trajectories collected with K baseline policies. Each of these policies can be quite suboptim…