1 paper · 1 filter
Alexandre Chenu, Nicolas Perrin-Gilbert, Stéphane Doncieux +1
Reinforcement learning agents need a reward signal to learn successful policies. When this signal is sparse or the corresponding gradient is deceptive, such agents need a dedicated…