works on

From the 1 of 20 linked papers with an AI index.

collaborators

20 papers

cs.LG2026

What's in a Smoothness Constant? Tighter Rates for Local SGD with Bounded Second-order Heterogeneity

Kumar Kshitij Patel, Rustem Islamov, Sebastian U Stich +3

The paper establishes tighter convergence rates for Local SGD (Federated Averaging) on general convex problems under a bounded second‑order heterogeneity assumption, and provides n…

cs.LG2026

Why Do We Need Warm-up? A Theoretical Perspective

Foivos Alimisis, Rustem Islamov, Aurelien Lucchi

Learning rate warm-up -- increasing the learning rate at the beginning of training -- has become a ubiquitous heuristic in modern deep learning, yet its theoretical foundations rem…

cs.LG2026

Beyond a Single Explanation of the Adam--SGD Gap

Chenxiang Zhang, Rustem Islamov, Enea Monzio Compagnoni +3

Prior work has identified several factors that can contribute to the performance gap between Adam and SGD, spanning data aspects, architecture design, and optimization properties.…

quant-ph2026

Gradient Scalability and Taylor Surrogation of Quantum Cost Landscapes

Sabri Meyer, Francesco Scala, Francesco Tacchino +1

Variational Quantum Algorithms are promising candidates for near-term quantum computing, yet they face scalability challenges due to barren plateaus, where gradients vanish exponen…

cs.LG2026

Byzantine-Robust and Differentially Private Federated Optimization under Weaker Assumptions

Rustem Islamov, Grigory Malinovsky, Alexander Gaponov +3

Federated Learning (FL) enables heterogeneous clients to collaboratively train a shared model without centralizing their raw data, offering an inherent level of privacy. However, g…

cs.LG2026

Where You Place the Norm Matters: From Prejudiced to Neutral Initializations

Emanuele Francazi, Francesco Pinto, Aurelien Lucchi +1

Normalization layers were introduced to stabilize and accelerate training, yet their influence is critical already at initialization, where they shape signal propagation and output…