9 citations · 17 across the 5 of their papers we have counts for
1 paper · 1 filter
Max Sobol Mark, Archit Sharma, Fahim Tajwar +3
It is desirable for policies to optimistically explore new states and behaviors during online reinforcement learning (RL) or fine-tuning, especially when prior offline data does no…