activity
20152023
most citedBootstrap your own latent: A new approach to self-supervised Learning

3.4k citations · 3.7k across the 29 of their papers we have counts for

collaborators

46 papers

math.PR2024

A New Bound on the Cumulant Generating Function of Dirichlet Processes

Pierre Perrault, Denis Belomestny, Pierre Ménard +4

In this paper, we introduce a novel approach for bounding the cumulant generating function (CGF) of a Dirichlet process (DP) , using superadditivity. In par…

cs.AI202314 cited

A General Theoretical Paradigm to Understand Learning from Human Preferences

Mohammad Gheshlaghi Azar, Mark Rowland, Bilal Piot +4

The prevalent deployment of learning from human preferences through reinforcement learning (RLHF) relies on two important approximations: the first assumes that pairwise preference…

cs.LG20221 cited

Understanding Self-Predictive Learning for Reinforcement Learning

Yunhao Tang, Zhaohan Daniel Guo, Pierre Harvey Richemond +13

We study the learning dynamics of self-predictive learning for reinforcement learning, a family of algorithms that learn representations by minimizing the prediction error of their…

stat.ML2022

Optimistic Posterior Sampling for Reinforcement Learning with Few Samples and Tight Guarantees

Daniil Tiapkin, Denis Belomestny, Daniele Calandriello +6

We consider reinforcement learning in an environment modeled by an episodic, finite, stage-dependent Markov decision process of horizon with states, and actions. The pe…

cs.LG2022

KL-Entropy-Regularized RL with a Generative Model is Minimax Optimal

Tadashi Kozuno, Wenhao Yang, Nino Vieillard +10

In this work, we consider and analyze the sample complexity of model-free reinforcement learning with a generative model. Particularly, we analyze mirror descent value iteration (M…

cs.LG2022

Marginalized Operators for Off-policy Reinforcement Learning

Yunhao Tang, Mark Rowland, Rémi Munos +1

In this work, we propose marginalized operators, a new class of off-policy evaluation operators for reinforcement learning. Marginalized operators strictly generalize generic multi…