collaborators

8 papers

cs.LG2026

Refined Analysis of Entropy-Regularized Actor-Critic

Safwan Labbi, Paul Mangold, Daniil Tiapkin +1

In this paper, we study the role of the critic in actor--critic for entropy-regularized, finite, discounted environments. We establish that, when the critic is exact, using the lat…

cs.LG2026

Uniform Diffusion Models Revisited: Leave-One-Out Denoiser and Absorbing State Reformulation

Samson Gourevitch, Yazid Janati, Dario Shariatian +4

Discrete diffusion models are often trained through clean-data prediction, but the prediction can be used in different ways to define the reverse dynamics. In Masked Diffusion Mode…

stat.ML2026

Gaussian Approximation and Multiplier Bootstrap for Federated Linear Stochastic Approximation

Ilya Levin, Maksim Shuklin, Eric Moulines +2

In this paper, we establish Berry-Esseen-type bounds for federated linear stochastic approximation (LSA). Our results provide the first federated Gaussian approximations for LSA th…

stat.ML2026

On Gaussian approximation for entropy-regularized Q-learning with function approximation

Artemy Rubtsov, Rahul Singh, Eric Moulines +2

In this paper, we derive rates of convergence in the high-dimensional central limit theorem for Polyak--Ruppert averaged iterates generated by entropy-regularized asynchronous Q-le…

cs.LG2026

Regret and Sample Complexity of Online Q-Learning via Concentration of Stochastic Approximation with Time-Inhomogeneous Markov Chains

Rahul Singh, Siddharth Chandak, Eric Moulines +2

We present the first regret bound for classical online Q-learning in infinite-horizon discounted Markov decision processes (MDPs), without relying on optimism or bonus terms. We fi…

cs.LG2026

Policy Gradient Methods for Non-Markovian Reinforcement Learning

Avik Kar, Siddharth Chandak, Rahul Singh +4

We study policy gradient methods for reinforcement learning in non-Markovian decision processes (NMDPs), where observations and rewards depend on the entire interaction history. To…