activity
20162022
most citedContrastive Behavioral Similarity Embeddings for Generalization in Reinforcement Learning

27 citations · 55 across the 5 of their papers we have counts for

collaborators

12 papers

cs.LG2022

Temporal Abstractions-Augmented Temporally Contrastive Learning: An Alternative to the Laplacian in RL

Akram Erraqabi, Marlos C. Machado, Mingde Zhao +4

In reinforcement learning, the graph Laplacian has proved to be a valuable tool in the task-agnostic setting, with applications ranging from skill discovery to reward shaping. Rece…

cs.LG202121 cited

On Bonus-Based Exploration Methods in the Arcade Learning Environment

Adrien Ali Taïga, William Fedus, Marlos C. Machado +2

Research on exploration in reinforcement learning, as applied to Atari 2600 game-playing, has emphasized tackling difficult exploration problems such as Montezuma's Revenge (Bellem…

cs.LG202127 cited

Contrastive Behavioral Similarity Embeddings for Generalization in Reinforcement Learning

Rishabh Agarwal, Marlos C. Machado, Pablo Samuel Castro +1

Reinforcement learning methods trained on few environments rarely learn policies that generalize to unseen environments. To improve generalization, we incorporate the inherent sequ…

cs.LG2020

Beyond variance reduction: Understanding the true impact of baselines on policy optimization

Wesley Chung, Valentin Thomas, Marlos C. Machado +1

Bandit and reinforcement learning (RL) problems can often be framed as optimization problems where the goal is to maximize average performance while having access only to stochasti…

cs.LG2020

An operator view of policy gradient methods

Dibya Ghosh, Marlos C. Machado, Nicolas Le Roux

We cast policy gradient methods as the repeated application of two operators: a policy improvement operator , which maps any policy to a better one ,…

cs.LG2018

Generalization and Regularization in DQN

Jesse Farebrother, Marlos C. Machado, Michael Bowling

Deep reinforcement learning algorithms have shown an impressive ability to learn complex control policies in high-dimensional tasks. However, despite the ever-increasing performanc…