works on

From the 1 of 18 linked papers with an AI index.

activity
20242026
collaborators

18 papers

cs.LG2026

What's in a Smoothness Constant? Tighter Rates for Local SGD with Bounded Second-order Heterogeneity

Kumar Kshitij Patel, Rustem Islamov, Sebastian U Stich +3

The paper establishes tighter convergence rates for Local SGD (Federated Averaging) on general convex problems under a bounded second‑order heterogeneity assumption, and provides n…

math.OC2026

The Dual Averaging Power-Prox Method with Application to Heavy-Tail Incremental Gradient

Yuan Gao, Jeremy Rack, Sebastian U. Stich

We study finite-sum composite optimization under two departures from classical stochastic gradient descent theory that are central in practice: incremental gradient access and heav…

cs.LG2026

Improved Convergence Analysis of Topology Dependence in Decentralized SGD

Yuki Takezawa, Anastasia Koloskova, Sebastian U. Stich

Decentralized SGD is a fundamental algorithm in decentralized learning, although the influence of an underlying network topology on its convergence behavior is not yet fully unders…

cs.LG2026

Forgetting Has Neighbors: Localized Collateral Forgetting in Machine Unlearning

Polina Dolgova, Sebastian U. Stich

Machine unlearning aims to remove the influence of selected training examples without full retraining. Standard evaluations often summarize unlearning quality with aggregate metric…

cs.LG2026

Enhancing LLM Training via Spectral Clipping

Xiaowen Jiang, Andrei Semenov, Sebastian U. Stich

While spectral-based optimizers like Muon operate directly on the spectrum of updates, standard adaptive methods such as AdamW do not account for the spectral structure of weights…

cs.LG2026

Learning When to Adapt

Ali Zindari, Xiaowen Jiang, Rotem Mulayoff +1

Low-rank adaptation (LoRA) is a widely used parameter-efficient fine-tuning method, yet its learned correction is static: the same low-rank update is applied to every input. This i…