11 citations · 47 across the 23 of their papers we have counts for
1 paper · 1 filter
Nirbhay Modhe, Qiaozi Gao, Ashwin Kalyan +3
Offline reinforcement learning (RL) methods strike a balance between exploration and exploitation by conservative value estimation -- penalizing values of unseen states and actions…