Publications (24)
Adaptive Compression for Communication-Efficient Distributed Training
Maksim Makarenko, Elnur Gasanov, Rustem Islamov +2
We propose Adaptive Compressed Gradient Descent (AdaCGD) - a novel optimization algorithm for communication-efficient training of supervised machine learning models with adaptive c…
Double Momentum and Error Feedback for Clipping with Fast Rates and Differential Privacy
Rustem Islamov, Samuel Horvath, Aurelien Lucchi +2
Strong Differential Privacy (DP) and Optimization guarantees are two desirable properties for a method in Federated Learning (FL). However, existing algorithms do not achieve both…
Loss Landscape Characterization of Neural Networks without Over-Parametrization
Rustem Islamov, Niccolò Ajroldi, Antonio Orvieto +1
Optimization methods play a crucial role in modern machine learning, powering the remarkable empirical achievements of deep learning models. These successes are even more remarkabl…
Adaptive Methods through the Lens of SDEs: Theoretical Insights on the Role of Noise
Enea Monzio Compagnoni, Tianlin Liu, Rustem Islamov +3
Despite the vast empirical evidence supporting the efficacy of adaptive optimization methods in deep learning, their theoretical understanding is far from complete. This work intro…
AsGrad: A Sharp Unified Analysis of Asynchronous-SGD Algorithms
Rustem Islamov, Mher Safaryan, Dan Alistarh
We analyze asynchronous-type algorithms for distributed SGD in the heterogeneous setting, where each worker has its own computation and communication speeds, as well as data distri…
Basis Matters: Better Communication-Efficient Second Order Methods for Federated Learning
Xun Qian, Rustem Islamov, Mher Safaryan +1
Recent advances in distributed optimization have shown that Newton-type methods with proper communication compression mechanisms can guarantee fast local rates and low communicatio…
EControl: Fast Distributed Optimization with Compression and Error Control
Yuan Gao, Rustem Islamov, Sebastian Stich
Modern distributed training relies heavily on communication compression to reduce the communication overhead. In this work, we study algorithms employing a popular class of contrac…
Byzantine-Robust and Differentially Private Federated Optimization under Weaker Assumptions
Rustem Islamov, Grigory Malinovsky, Alexander Gaponov +3
Federated Learning (FL) enables heterogeneous clients to collaboratively train a shared model without centralizing their raw data, offering an inherent level of privacy. However, g…
What's in a Smoothness Constant? Tighter Rates for Local SGD with Bounded Second-order Heterogeneity
Kumar Kshitij Patel, Rustem Islamov, Sebastian U Stich +3
The paper establishes tighter convergence rates for Local SGD (Federated Averaging) on general convex problems under a bounded second‑order heterogeneity assumption, and provides n…
Why Do We Need Warm-up? A Theoretical Perspective
Foivos Alimisis, Rustem Islamov, Aurelien Lucchi
Learning rate warm-up -- increasing the learning rate at the beginning of training -- has become a ubiquitous heuristic in modern deep learning, yet its theoretical foundations rem…
Adaptive Methods Are Preferable in High Privacy Settings: An SDE Perspective
Enea Monzio Compagnoni, Alessandro Stanghellini, Rustem Islamov +2
Differential Privacy (DP) is becoming central to large-scale training as privacy regulations tighten. We revisit how DP noise interacts with adaptivity in optimization through the…
Distributed Second Order Methods with Fast Rates and Compressed Communication
Rustem Islamov, Xun Qian, Peter Richtárik
We develop several new communication-efficient second-order methods for distributed optimization. Our first method, NEWTON-STAR, is a variant of Newton's method from which it inher…
Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size
Rustem Islamov, Niccolo Ajroldi, Antonio Orvieto +1
Modern optimization algorithms that incorporate momentum and adaptive step-size offer improved performance in numerous challenging deep learning tasks. However, their effectiveness…
Unbiased and Sign Compression in Distributed Learning: Comparing Noise Resilience via SDEs
Enea Monzio Compagnoni, Rustem Islamov, Frank Norbert Proske +1
Distributed methods are essential for handling machine learning pipelines comprising large-scale models and datasets. However, their benefits often come at the cost of increased co…
FedNL: Making Newton-Type Methods Applicable to Federated Learning
Mher Safaryan, Rustem Islamov, Xun Qian +1
Inspired by recent work of Islamov et al (2021), we propose a family of Federated Newton Learn (FedNL) methods, which we believe is a marked step in the direction of making second-…
Non-Euclidean Gradient Descent Operates at the Edge of Stability
Rustem Islamov, Michael Crawshaw, Jeremy Cohen +1
The Edge of Stability (EoS) is a phenomenon where the sharpness (largest eigenvalue) of the Hessian approaches and then hovers near the stability threshold during gradient d…
On the Role of Batch Size in Stochastic Conditional Gradient Methods
Rustem Islamov, Roman Machacek, Aurelien Lucchi +3
We study the role of batch size in stochastic conditional gradient methods under a -Kurdyka-Åojasiewicz (-KL) condition. Focusing on momentum-based stochastic conditional…
Distributed Newton-Type Methods with Communication Compression and Bernoulli Aggregation
Rustem Islamov, Xun Qian, SlavomÃr Hanzely +2
Despite their high computation and communication costs, Newton-type methods remain an appealing option for distributed training due to their robustness against ill-conditioned conv…
Towards Faster Decentralized Stochastic Optimization with Communication Compression
Rustem Islamov, Yuan Gao, Sebastian U. Stich
Communication efficiency has garnered significant attention as it is considered the main bottleneck for large-scale decentralized Machine Learning applications in distributed and f…
Safe-EF: Error Feedback for Nonsmooth Constrained Optimization
Rustem Islamov, Yarden As, Ilyas Fatkhullin
Federated learning faces severe communication bottlenecks due to the high dimensionality of model updates. Communication compression with contractive compressors (e.g., Top-K) is o…
Clip21: Error Feedback for Gradient Clipping
Sarit Khirirat, Eduard Gorbunov, Samuel Horváth +3
Motivated by the increasing popularity and importance of large-scale training under differential privacy (DP) constraints, we study distributed gradient methods with gradient clipp…
On the Interaction of Batch Noise, Adaptivity, and Compression, under -Smoothness: An SDE Approach
Enea Monzio Compagnoni, Rustem Islamov, Frank Norbert Proske +3
Distributed stochastic optimization intertwines (i) stochastic gradient noise, (ii) communication compression, and (iii) adaptive/normalized updates. While each factor has been stu…
Beyond a Single Explanation of the Adam--SGD Gap
Chenxiang Zhang, Rustem Islamov, Enea Monzio Compagnoni +3
Prior work has identified several factors that can contribute to the performance gap between Adam and SGD, spanning data aspects, architecture design, and optimization properties.…
Partially Personalized Federated Learning: Breaking the Curse of Data Heterogeneity
Konstantin Mishchenko, Rustem Islamov, Eduard Gorbunov +1
We present a partially personalized formulation of Federated Learning (FL) that strikes a balance between the flexibility of personalization and cooperativeness of global training.…