Showing stat.MLShow all
3 papers · 1 filter
stat.ML2025
Model-free Posterior Sampling via Learning Rate Randomization
Daniil Tiapkin, Denis Belomestny, Daniele Calandriello +6
In this paper, we introduce Randomized Q-learning (RandQL), a novel randomized model-free algorithm for regret minimization in episodic Markov Decision Processes (MDPs). To the bes…
stat.ML2025
Forward Reverse Kernel Regression for the Schrödinger bridge problem
Denis Belomestny, John. Schoenmakers
In this paper, we study the Schrödinger Bridge Problem (SBP), which is central to entropic optimal transport. For general reference processes and begin--endpoint distributions, we…
stat.ML2024
Demonstration-Regularized RL
Daniil Tiapkin, Denis Belomestny, Daniele Calandriello +5
Incorporating expert demonstrations has empirically helped to improve the sample efficiency of reinforcement learning (RL). This paper quantifies theoretically to what extent this…