4 papers
Improved Convergence Analysis of Topology Dependence in Decentralized SGD
Yuki Takezawa, Anastasia Koloskova, Sebastian U. Stich
Decentralized SGD is a fundamental algorithm in decentralized learning, although the influence of an underlying network topology on its convergence behavior is not yet fully unders…
FedMuon: Federated Learning with Bias-corrected LMO-based Optimization
Yuki Takezawa, Anastasia Koloskova, Xiaowen Jiang +1
Recently, a new optimization method based on the linear minimization oracle (LMO), called Muon, has been attracting increasing attention since it can train neural networks faster t…
Exploiting Similarity for Computation and Communication-Efficient Decentralized Optimization
Yuki Takezawa, Xiaowen Jiang, Anton Rodomanov +1
Reducing communication complexity is critical for efficient decentralized optimization. The proximal decentralized optimization (PDO) framework is particularly appealing, as method…
Scalable Decentralized Learning with Teleportation
Yuki Takezawa, Sebastian U. Stich
Decentralized SGD can run with low communication costs, but its sparse communication characteristics deteriorate the convergence rate, especially when the number of nodes is large.…