3.4k citations · 4k across the 14 of their papers we have counts for
4 papers · 1 filter
A General Theoretical Paradigm to Understand Learning from Human Preferences
Mohammad Gheshlaghi Azar, Mark Rowland, Bilal Piot +4
The prevalent deployment of learning from human preferences through reinforcement learning (RLHF) relies on two important approximations: the first assumes that pairwise preference…
The Advantage Regret-Matching Actor-Critic
Audrūnas Gruslys, Marc Lanctot, Rémi Munos +10
Regret minimization has played a key role in online learning, equilibrium computation in games, and reinforcement learning (RL). In this paper, we describe a general model-free RL…
World Discovery Models
Mohammad Gheshlaghi Azar, Bilal Piot, Bernardo Avila Pires +3
As humans we are driven by a strong desire for seeking novelty in our world. Also upon observing a novel pattern we are capable of refining our understanding of the world based on…
Rainbow: Combining Improvements in Deep Reinforcement Learning
Matteo Hessel, Joseph Modayil, Hado van Hasselt +7
The deep reinforcement learning community has made several independent improvements to the DQN algorithm. However, it is unclear which of these extensions are complementary and can…