2 papers
cs.LG2026
Rethinking Bregman Divergences in Kronecker-Factored Optimizers
Bing Liu, Wenjie Zhou, Chengcheng Zhao
Shampoo-style optimizers approximate gradient covariance matrices using Kronecker-factored structures. Recent work~\cite{lin2026understanding} showed that such approximations can b…
cs.LG2026
Row-Stochastic Matrices Can Provably Outperform Doubly Stochastic Matrices in Decentralized Learning
Bing Liu, Boao Kong, Limin Lu +2
Decentralized learning often involves a weighted global loss with heterogeneous node weights . We revisit two natural strategies for incorporating these weights: (i) embedding t…