6 citations · 10 across the 5 of their papers we have counts for
4 papers · 1 filter
A single algorithm for both restless and rested rotting bandits
Julien Seznec, Pierre Ménard, Alessandro Lazaric +1
In many application domains (e.g., recommender systems, intelligent tutoring systems), the rewards associated to the actions tend to decrease over time. This decay is either caused…
Proximal Point Nash Learning from Human Feedback
Daniil Tiapkin, Daniele Calandriello, Denis Belomestny +5
Traditional Reinforcement Learning from Human Feedback (RLHF) often relies on reward models, frequently assuming preference structures like the Bradley--Terry model, which may not…
Model-free Posterior Sampling via Learning Rate Randomization
Daniil Tiapkin, Denis Belomestny, Daniele Calandriello +6
In this paper, we introduce Randomized Q-learning (RandQL), a novel randomized model-free algorithm for regret minimization in episodic Markov Decision Processes (MDPs). To the bes…
Demonstration-Regularized RL
Daniil Tiapkin, Denis Belomestny, Daniele Calandriello +5
Incorporating expert demonstrations has empirically helped to improve the sample efficiency of reinforcement learning (RL). This paper quantifies theoretically to what extent this…