3 citations · 3 across the 3 of their papers we have counts for
1 paper · 1 filter
Sungryull Sohn, Yinlam Chow, Jayden Ooi +4
In batch reinforcement learning (RL), one often constrains a learned policy to be close to the behavior (data-generating) policy, e.g., by constraining the learned action distribut…