91 citations · 103 across the 2 of their papers we have counts for
1 paper · 1 filter
Evan Shelhamer, Parsa Mahmoudieh, Max Argus +1
Reinforcement learning optimizes policies for expected cumulative reward. Need the supervision be so narrow? Reward is delayed and sparse for many tasks, making it a difficult and…