collaborators

14 papers

cs.LG2026

Proximal Policy Optimization for Amortized Discrete Sampling

Anna Zykova-Myzina, Timofei Gritsaev, Daniil Tiapkin +1

This paper explores policy gradient algorithms for training stochastic policies to sample from structured discrete probability distributions under the Generative Flow Network (GFlo…

cs.LG2026

Refined Analysis of Entropy-Regularized Actor-Critic

Safwan Labbi, Paul Mangold, Daniil Tiapkin +1

In this paper, we study the role of the critic in actor--critic for entropy-regularized, finite, discounted environments. We establish that, when the critic is exact, using the lat…

cs.LG2026

Beyond Softmax and Entropy: Convergence Rates of Policy Gradients with f-SoftArgmax Parameterization & Coupled Regularization

Safwan Labbi, Daniil Tiapkin, Paul Mangold +1

Policy gradient methods are known to be highly sensitive to the choice of policy parameterization. In particular, the widely used softmax parameterization can induce ill-conditione…

cs.LG2026

On Global Convergence Rates for Federated Softmax Policy Gradient under Heterogeneous Environments

Safwan Labbi, Paul Mangold, Daniil Tiapkin +1

We provide global convergence rates for vanilla and entropy-regularized federated softmax stochastic policy gradient (FedPG) with local training. We show that FedPG converges to a…

stat.ML2026

Proximal Point Nash Learning from Human Feedback

Daniil Tiapkin, Daniele Calandriello, Denis Belomestny +5

Traditional Reinforcement Learning from Human Feedback (RLHF) often relies on reward models, frequently assuming preference structures like the Bradley--Terry model, which may not…

cs.LG2026

Learning Shortest Paths with Generative Flow Networks

Nikita Morozov, Ian Maksimov, Daniil Tiapkin +1

In this paper, we present a novel learning framework for finding shortest paths in graphs utilizing Generative Flow Networks (GFlowNets). First, we examine theoretical properties o…