19 papers
On Gaussian approximation for entropy-regularized Q-learning with function approximation
Artemy Rubtsov, Rahul Singh, Eric Moulines +2
In this paper, we derive rates of convergence in the high-dimensional central limit theorem for Polyak--Ruppert averaged iterates generated by entropy-regularized asynchronous Q-le…
Gaussian Approximation and Multiplier Bootstrap for Stochastic Gradient Descent
Marina Sheshukova, Sergey Samsonov, Denis Belomestny +4
In this paper, we establish the non-asymptotic validity of the multiplier bootstrap procedure for constructing the confidence sets using the Stochastic Gradient Descent (SGD) algor…
Gaussian Approximation for Asynchronous Q-learning
Artemy Rubtsov, Sergey Samsonov, Vladimir Ulyanov +1
In this paper, we derive rates of convergence in the high-dimensional central limit theorem for Polyak-Ruppert averaged iterates generated by the asynchronous Q-learning algorithm…
Proximal Point Nash Learning from Human Feedback
Daniil Tiapkin, Daniele Calandriello, Denis Belomestny +5
Traditional Reinforcement Learning from Human Feedback (RLHF) often relies on reward models, frequently assuming preference structures like the Bradley--Terry model, which may not…
Tight Bounds for Schrödinger Potential Estimation in Unpaired Data Translation
Nikita Puchkin, Denis Suchkov, Alexey Naumov +1
Modern methods of generative modelling and unpaired data translation based on Schrödinger bridges and stochastic optimal control theory aim to transform an initial density to a ta…
Schrödinger bridge problem via empirical risk minimization
Denis Belomestny, Alexey Naumov, Nikita Puchkin +1
We study the Schrödinger bridge problem when the endpoint distributions are available only through samples. Classical computational approaches estimate Schrödinger potentials via…