5 papers · 1 filter
Proximal Point Nash Learning from Human Feedback
Daniil Tiapkin, Daniele Calandriello, Denis Belomestny +5
Traditional Reinforcement Learning from Human Feedback (RLHF) often relies on reward models, frequently assuming preference structures like the Bradley--Terry model, which may not…
Model-free Posterior Sampling via Learning Rate Randomization
Daniil Tiapkin, Denis Belomestny, Daniele Calandriello +6
In this paper, we introduce Randomized Q-learning (RandQL), a novel randomized model-free algorithm for regret minimization in episodic Markov Decision Processes (MDPs). To the bes…
Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games
Antonio Ocello, Daniil Tiapkin, Lorenzo Mancini +2
We introduce Mean-Field Trust Region Policy Optimization (MF-TRPO), a novel algorithm designed to compute approximate Nash equilibria for ergodic Mean-Field Games (MFG) in finite s…
Improved High-Probability Bounds for the Temporal Difference Learning Algorithm via Exponential Stability
Sergey Samsonov, Daniil Tiapkin, Alexey Naumov +1
In this paper we consider the problem of obtaining sharp bounds for the performance of temporal difference (TD) methods with linear function approximation for policy evaluation in…
Demonstration-Regularized RL
Daniil Tiapkin, Denis Belomestny, Daniele Calandriello +5
Incorporating expert demonstrations has empirically helped to improve the sample efficiency of reinforcement learning (RL). This paper quantifies theoretically to what extent this…