activity
20172022
most citedThe Cramer Distance as a Solution to Biased Wasserstein Gradients

256 citations · 1.1k across the 20 of their papers we have counts for

collaborators

31 papers

cs.LG202121 cited

On Bonus-Based Exploration Methods in the Arcade Learning Environment

Adrien Ali Taïga, William Fedus, Marlos C. Machado +2

Research on exploration in reinforcement learning, as applied to Atari 2600 game-playing, has emphasized tackling difficult exploration problems such as Montezuma's Revenge (Bellem…

cs.LG2021

Metrics and continuity in reinforcement learning

Charline Le Lan, Marc G. Bellemare, Pablo Samuel Castro

In most practical applications of reinforcement learning, it is untenable to maintain direct estimates for individual states; in continuous-state systems, it is impossible. Instead…

cs.LG202127 cited

Contrastive Behavioral Similarity Embeddings for Generalization in Reinforcement Learning

Rishabh Agarwal, Marlos C. Machado, Pablo Samuel Castro +1

Reinforcement learning methods trained on few environments rarely learn policies that generalize to unseen environments. To improve generalization, we incorporate the inherent sequ…

cs.AI2020

The Importance of Pessimism in Fixed-Dataset Policy Optimization

Jacob Buckman, Carles Gelada, Marc G. Bellemare

We study worst-case guarantees on the expected return of fixed-dataset policy optimization algorithms. Our core contribution is a unified conceptual and mathematical framework for…

cs.LG20208 cited

Representations for Stable Off-Policy Reinforcement Learning

Dibya Ghosh, Marc G. Bellemare

Reinforcement learning with function approximation can be unstable and even divergent, especially when combined with off-policy learning and Bellman updates. In deep reinforcement…

cs.LG2020

The Value-Improvement Path: Towards Better Representations for Reinforcement Learning

Will Dabney, André Barreto, Mark Rowland +4

In value-based reinforcement learning (RL), unlike in supervised learning, the agent faces not a single, stationary, approximation problem, but a sequence of value prediction probl…