From the 1 of 11 linked papers with an AI index.
11 papers
What's in a Smoothness Constant? Tighter Rates for Local SGD with Bounded Second-order Heterogeneity
Kumar Kshitij Patel, Rustem Islamov, Sebastian U Stich +3
The paper establishes tighter convergence rates for Local SGD (Federated Averaging) on general convex problems under a bounded second‑order heterogeneity assumption, and provides n…
Non-Euclidean Gradient Descent Operates at the Edge of Stability
Rustem Islamov, Michael Crawshaw, Jeremy Cohen +1
The Edge of Stability (EoS) is a phenomenon where the sharpness (largest eigenvalue) of the Hessian approaches and then hovers near the stability threshold during gradient d…
Why Do We Need Warm-up? A Theoretical Perspective
Foivos Alimisis, Rustem Islamov, Aurelien Lucchi
Learning rate warm-up -- increasing the learning rate at the beginning of training -- has become a ubiquitous heuristic in modern deep learning, yet its theoretical foundations rem…
Beyond a Single Explanation of the Adam--SGD Gap
Chenxiang Zhang, Rustem Islamov, Enea Monzio Compagnoni +3
Prior work has identified several factors that can contribute to the performance gap between Adam and SGD, spanning data aspects, architecture design, and optimization properties.…
On the Interaction of Batch Noise, Adaptivity, and Compression, under -Smoothness: An SDE Approach
Enea Monzio Compagnoni, Rustem Islamov, Frank Norbert Proske +3
Distributed stochastic optimization intertwines (i) stochastic gradient noise, (ii) communication compression, and (iii) adaptive/normalized updates. While each factor has been stu…
Byzantine-Robust and Differentially Private Federated Optimization under Weaker Assumptions
Rustem Islamov, Grigory Malinovsky, Alexander Gaponov +3
Federated Learning (FL) enables heterogeneous clients to collaboratively train a shared model without centralizing their raw data, offering an inherent level of privacy. However, g…