14 papers
Proximal Policy Optimization for Amortized Discrete Sampling
Anna Zykova-Myzina, Timofei Gritsaev, Daniil Tiapkin +1
This paper explores policy gradient algorithms for training stochastic policies to sample from structured discrete probability distributions under the Generative Flow Network (GFlo…
Refined Analysis of Entropy-Regularized Actor-Critic
Safwan Labbi, Paul Mangold, Daniil Tiapkin +1
In this paper, we study the role of the critic in actor--critic for entropy-regularized, finite, discounted environments. We establish that, when the critic is exact, using the lat…
Beyond Softmax and Entropy: Convergence Rates of Policy Gradients with f-SoftArgmax Parameterization & Coupled Regularization
Safwan Labbi, Daniil Tiapkin, Paul Mangold +1
Policy gradient methods are known to be highly sensitive to the choice of policy parameterization. In particular, the widely used softmax parameterization can induce ill-conditione…
On Global Convergence Rates for Federated Softmax Policy Gradient under Heterogeneous Environments
Safwan Labbi, Paul Mangold, Daniil Tiapkin +1
We provide global convergence rates for vanilla and entropy-regularized federated softmax stochastic policy gradient (FedPG) with local training. We show that FedPG converges to a…
Proximal Point Nash Learning from Human Feedback
Daniil Tiapkin, Daniele Calandriello, Denis Belomestny +5
Traditional Reinforcement Learning from Human Feedback (RLHF) often relies on reward models, frequently assuming preference structures like the Bradley--Terry model, which may not…
Learning Shortest Paths with Generative Flow Networks
Nikita Morozov, Ian Maksimov, Daniil Tiapkin +1
In this paper, we present a novel learning framework for finding shortest paths in graphs utilizing Generative Flow Networks (GFlowNets). First, we examine theoretical properties o…