activity
20152026
most citedBootstrap your own latent: A new approach to self-supervised Learning

3.4k citations · 4.5k across the 104 of their papers we have counts for

collaborators
Showing 2023Show all

12 papers · 1 filter

stat.ML2023★ 1 cited

Nash Learning from Human Feedback

Rémi Munos, Michal Valko, Daniele Calandriello +14

Reinforcement learning from human feedback (RLHF) has emerged as the main paradigm for aligning large language models (LLMs) with human preferences. Typically, RLHF involves the in…

cs.AI2023★ 14 cited

A General Theoretical Paradigm to Understand Learning from Human Preferences

Mohammad Gheshlaghi Azar, Mark Rowland, Bilal Piot +4

The prevalent deployment of learning from human preferences through reinforcement learning (RLHF) relies on two important approximations: the first assumes that pairwise preference…

stat.ML2023

Model-free Posterior Sampling via Learning Rate Randomization

Daniil Tiapkin, Denis Belomestny, Daniele Calandriello +6

In this paper, we introduce Randomized Q-learning (RandQL), a novel randomized model-free algorithm for regret minimization in episodic Markov Decision Processes (MDPs). To the bes…

stat.ML2023

Demonstration-Regularized RL

Daniil Tiapkin, Denis Belomestny, Daniele Calandriello +5

Incorporating expert demonstrations has empirically helped to improve the sample efficiency of reinforcement learning (RL). This paper quantifies theoretically to what extent this…

cs.GT2023

Local and adaptive mirror descents in extensive-form games

Côme Fiegel, Pierre Ménard, Tadashi Kozuno +3

We study how to learn -optimal strategies in zero-sum imperfect information games (IIG) with trajectory feedback. In this setting, players update their policies sequentially bas…

cs.LG2023

Half-Hop: A graph upsampling approach for slowing down message passing

Mehdi Azabou, Venkataramana Ganesh, Shantanu Thakoor +6

Message passing neural networks have shown a lot of success on graph-structured data. However, there are many instances where message passing can lead to over-smoothing or fail whe…