From the 1 of 18 linked papers with an AI index.
18 papers
What's in a Smoothness Constant? Tighter Rates for Local SGD with Bounded Second-order Heterogeneity
Kumar Kshitij Patel, Rustem Islamov, Sebastian U Stich +3
The paper establishes tighter convergence rates for Local SGD (Federated Averaging) on general convex problems under a bounded second‑order heterogeneity assumption, and provides n…
The Dual Averaging Power-Prox Method with Application to Heavy-Tail Incremental Gradient
Yuan Gao, Jeremy Rack, Sebastian U. Stich
We study finite-sum composite optimization under two departures from classical stochastic gradient descent theory that are central in practice: incremental gradient access and heav…
Improved Convergence Analysis of Topology Dependence in Decentralized SGD
Yuki Takezawa, Anastasia Koloskova, Sebastian U. Stich
Decentralized SGD is a fundamental algorithm in decentralized learning, although the influence of an underlying network topology on its convergence behavior is not yet fully unders…
Forgetting Has Neighbors: Localized Collateral Forgetting in Machine Unlearning
Polina Dolgova, Sebastian U. Stich
Machine unlearning aims to remove the influence of selected training examples without full retraining. Standard evaluations often summarize unlearning quality with aggregate metric…
Enhancing LLM Training via Spectral Clipping
Xiaowen Jiang, Andrei Semenov, Sebastian U. Stich
While spectral-based optimizers like Muon operate directly on the spectrum of updates, standard adaptive methods such as AdamW do not account for the spectral structure of weights…
Learning When to Adapt
Ali Zindari, Xiaowen Jiang, Rotem Mulayoff +1
Low-rank adaptation (LoRA) is a widely used parameter-efficient fine-tuning method, yet its learned correction is static: the same low-rank update is applied to every input. This i…