5 papers
Gradient Descent as a Perceptron Algorithm: Understanding Dynamics and Implicit Acceleration
Alexander Tyurin
Even for the gradient descent (GD) method applied to neural network training, understanding its optimization dynamics, including convergence rate, iterate trajectories, function va…
Asynchronous Policy Gradient Aggregation for Efficient Distributed Reinforcement Learning
Alexander Tyurin, Andrei Spiridonov, Varvara Rudenko
We study distributed reinforcement learning (RL) with policy gradient methods under asynchronous and parallel computations and communications. While non-distributed methods are wel…
Proving the Limited Scalability of Centralized Distributed Optimization via a New Lower Bound Construction
Alexander Tyurin
We consider centralized distributed optimization in the classical federated learning setup, where workers jointly find an -stationary point of an -smooth, -d…
Birch SGD: A Tree Graph Framework for Local and Asynchronous SGD Methods
Alexander Tyurin, Danil Sivtsov
We propose a new unifying framework, Birch SGD, for analyzing and designing distributed SGD methods. The central idea is to represent each method as a weighted directed tree, refer…
Learning of Population Dynamics: Inverse Optimization Meets JKO Scheme
Mikhail Persiianov, Jiawei Chen, Petr Mokrov +3
Learning population dynamics involves recovering the underlying process that governs particle evolution, given evolutionary snapshots of samples at discrete time points. Recent met…