91 citations · 91 across the 1 of their papers we have counts for
1 paper
Evan Shelhamer, Parsa Mahmoudieh, Max Argus +1
Reinforcement learning optimizes policies for expected cumulative reward. Need the supervision be so narrow? Reward is delayed and sparse for many tasks, making it a difficult and…