256 citations · 1.1k across the 20 of their papers we have counts for
31 papers
On Bonus-Based Exploration Methods in the Arcade Learning Environment
Adrien Ali Taïga, William Fedus, Marlos C. Machado +2
Research on exploration in reinforcement learning, as applied to Atari 2600 game-playing, has emphasized tackling difficult exploration problems such as Montezuma's Revenge (Bellem…
Metrics and continuity in reinforcement learning
Charline Le Lan, Marc G. Bellemare, Pablo Samuel Castro
In most practical applications of reinforcement learning, it is untenable to maintain direct estimates for individual states; in continuous-state systems, it is impossible. Instead…
Contrastive Behavioral Similarity Embeddings for Generalization in Reinforcement Learning
Rishabh Agarwal, Marlos C. Machado, Pablo Samuel Castro +1
Reinforcement learning methods trained on few environments rarely learn policies that generalize to unseen environments. To improve generalization, we incorporate the inherent sequ…
The Importance of Pessimism in Fixed-Dataset Policy Optimization
Jacob Buckman, Carles Gelada, Marc G. Bellemare
We study worst-case guarantees on the expected return of fixed-dataset policy optimization algorithms. Our core contribution is a unified conceptual and mathematical framework for…
Representations for Stable Off-Policy Reinforcement Learning
Dibya Ghosh, Marc G. Bellemare
Reinforcement learning with function approximation can be unstable and even divergent, especially when combined with off-policy learning and Bellman updates. In deep reinforcement…
The Value-Improvement Path: Towards Better Representations for Reinforcement Learning
Will Dabney, André Barreto, Mark Rowland +4
In value-based reinforcement learning (RL), unlike in supervised learning, the agent faces not a single, stationary, approximation problem, but a sequence of value prediction probl…