13 citations · 28 across the 3 of their papers we have counts for
1 paper · 1 filter
Edoardo Cetin, Philip J. Ball, Steve Roberts +1
Off-policy reinforcement learning (RL) from pixel observations is notoriously unstable. As a result, many successful algorithms must combine different domain-specific practices and…