8 papers
Refined Analysis of Entropy-Regularized Actor-Critic
Safwan Labbi, Paul Mangold, Daniil Tiapkin +1
In this paper, we study the role of the critic in actor--critic for entropy-regularized, finite, discounted environments. We establish that, when the critic is exact, using the lat…
Uniform Diffusion Models Revisited: Leave-One-Out Denoiser and Absorbing State Reformulation
Samson Gourevitch, Yazid Janati, Dario Shariatian +4
Discrete diffusion models are often trained through clean-data prediction, but the prediction can be used in different ways to define the reverse dynamics. In Masked Diffusion Mode…
Gaussian Approximation and Multiplier Bootstrap for Federated Linear Stochastic Approximation
Ilya Levin, Maksim Shuklin, Eric Moulines +2
In this paper, we establish Berry-Esseen-type bounds for federated linear stochastic approximation (LSA). Our results provide the first federated Gaussian approximations for LSA th…
On Gaussian approximation for entropy-regularized Q-learning with function approximation
Artemy Rubtsov, Rahul Singh, Eric Moulines +2
In this paper, we derive rates of convergence in the high-dimensional central limit theorem for Polyak--Ruppert averaged iterates generated by entropy-regularized asynchronous Q-le…
Regret and Sample Complexity of Online Q-Learning via Concentration of Stochastic Approximation with Time-Inhomogeneous Markov Chains
Rahul Singh, Siddharth Chandak, Eric Moulines +2
We present the first regret bound for classical online Q-learning in infinite-horizon discounted Markov decision processes (MDPs), without relying on optimism or bonus terms. We fi…
Policy Gradient Methods for Non-Markovian Reinforcement Learning
Avik Kar, Siddharth Chandak, Rahul Singh +4
We study policy gradient methods for reinforcement learning in non-Markovian decision processes (NMDPs), where observations and rewards depend on the entire interaction history. To…