5 papers
Adaptive Regularized Newton Method with Inexact Hessian
Aleksandr Shestakov, Nail Bashirov, Andrei Semenov +4
Newton's method is the most widespread high-order method, demanding the gradient and the Hessian of the objective function. However, one of the main disadvantages of Newtons method…
Benchmarking Optimizers for Large Language Model Pretraining
Andrei Semenov, Matteo Pagliardini, Martin Jaggi
The recent development of Large Language Models (LLMs) has been accompanied by an effervescence of novel ideas and methods to better optimize the loss of deep learning models. Clai…
Gradient-Normalized Smoothness for Optimization with Approximate Hessians
Andrei Semenov, Martin Jaggi, Nikita Doikov
In this work, we develop new optimization algorithms that use approximate second-order information combined with the gradient regularization technique to achieve fast global conver…
Sign Operator for Coping with Heavy-Tailed Noise in Non-Convex Optimization: High Probability Bounds Under -Smoothness
Nikita Kornilov, Philip Zmushko, Andrei Semenov +3
In recent years, non-convex optimization problems are more often described by generalized -smoothness assumption rather than standard one. Meanwhile, severely corrupted…
Just a Simple Transformation is Enough for Data Protection in Vertical Federated Learning
Andrei Semenov, Philip Zmushko, Alexander Pichugin +1
Vertical Federated Learning (VFL) aims to enable collaborative training of deep learning models while maintaining privacy protection. However, the VFL procedure still has component…