3.4k citations · 4k across the 12 of their papers we have counts for
14 papers
A General Theoretical Paradigm to Understand Learning from Human Preferences
Mohammad Gheshlaghi Azar, Mark Rowland, Bilal Piot +4
The prevalent deployment of learning from human preferences through reinforcement learning (RLHF) relies on two important approximations: the first assumes that pairwise preference…
Understanding Self-Predictive Learning for Reinforcement Learning
Yunhao Tang, Zhaohan Daniel Guo, Pierre Harvey Richemond +13
We study the learning dynamics of self-predictive learning for reinforcement learning, a family of algorithms that learn representations by minimizing the prediction error of their…
KL-Entropy-Regularized RL with a Generative Model is Minimax Optimal
Tadashi Kozuno, Wenhao Yang, Nino Vieillard +10
In this work, we consider and analyze the sample complexity of model-free reinforcement learning with a generative model. Particularly, we analyze mirror descent value iteration (M…
Drop, Swap, and Generate: A Self-Supervised Approach for Generating Neural Activity
Ran Liu, Mehdi Azabou, Max Dabagia +5
Meaningful and simplified representations of neural activity can yield insights into how and what information is being processed within a neural circuit. However, without labels, f…
Geometric Entropic Exploration
Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, Alaa Saade +7
Exploration is essential for solving complex Reinforcement Learning (RL) tasks. Maximum State-Visitation Entropy (MSVE) formulates the exploration problem as a well-defined policy…
The Advantage Regret-Matching Actor-Critic
Audrūnas Gruslys, Marc Lanctot, Rémi Munos +10
Regret minimization has played a key role in online learning, equilibrium computation in games, and reinforcement learning (RL). In this paper, we describe a general model-free RL…