1 citations · 1 across the 1 of their papers we have counts for
1 paper · 1 filter
Muhan Lin, Shuyang Shi, Yue Guo +6
The correct specification of reward models is a well-known challenge in reinforcement learning. Hand-crafted reward functions often lead to inefficient or suboptimal policies and m…