activity
20162026
most citedSecond-Order Kernel Online Convex Optimization with Adaptive Sketching

23 citations · 112 across the 33 of their papers we have counts for

collaborators
Showing 2023Show all

6 papers · 1 filter

stat.ML2023★ 1 cited

Nash Learning from Human Feedback

Rémi Munos, Michal Valko, Daniele Calandriello +14

Reinforcement learning from human feedback (RLHF) has emerged as the main paradigm for aligning large language models (LLMs) with human preferences. Typically, RLHF involves the in…

cs.AI2023★ 14 cited

A General Theoretical Paradigm to Understand Learning from Human Preferences

Mohammad Gheshlaghi Azar, Mark Rowland, Bilal Piot +4

The prevalent deployment of learning from human preferences through reinforcement learning (RLHF) relies on two important approximations: the first assumes that pairwise preference…

stat.ML2023

Model-free Posterior Sampling via Learning Rate Randomization

Daniil Tiapkin, Denis Belomestny, Daniele Calandriello +6

In this paper, we introduce Randomized Q-learning (RandQL), a novel randomized model-free algorithm for regret minimization in episodic Markov Decision Processes (MDPs). To the bes…

stat.ML2023

Demonstration-Regularized RL

Daniil Tiapkin, Denis Belomestny, Daniele Calandriello +5

Incorporating expert demonstrations has empirically helped to improve the sample efficiency of reinforcement learning (RL). This paper quantifies theoretically to what extent this…

cs.LG2023

Unlocking the Power of Representations in Long-term Novelty-based Exploration

Alaa Saade, Steven Kapturowski, Daniele Calandriello +6

We introduce Robust Exploration via Clustering-based Online Density Estimation (RECODE), a non-parametric method for novelty-based exploration that estimates visitation counts for…

stat.ML2023

Fast Rates for Maximum Entropy Exploration

Daniil Tiapkin, Denis Belomestny, Daniele Calandriello +7

We address the challenge of exploration in reinforcement learning (RL) when the agent operates in an unknown environment with sparse or no rewards. In this work, we study the maxim…