4 papers
Communication-Efficient Gluon in Federated Learning
Xun Qian, Alexander Gaponov, Grigory Malinovsky +1
Recent developments have shown that Muon-type optimizers based on linear minimization oracles (LMOs) over non-Euclidean norm balls have the potential to get superior practical perf…
Byzantine-Robust and Differentially Private Federated Optimization under Weaker Assumptions
Rustem Islamov, Grigory Malinovsky, Alexander Gaponov +3
Federated Learning (FL) enables heterogeneous clients to collaboratively train a shared model without centralizing their raw data, offering an inherent level of privacy. However, g…
Error Feedback for Muon and Friends
Kaja Gruntkowska, Alexander Gaponov, Zhirayr Tovmasyan +1
Recent optimizers like Muon, Scion, and Gluon have pushed the frontier of large-scale deep learning by exploiting layer-wise linear minimization oracles (LMOs) over non-Euclidean n…
Pay Attention to Attention Distribution: A New Local Lipschitz Bound for Transformers
Nikolay Yudin, Alexander Gaponov, Sergei Kudriashov +1
We introduce a novel upper bound on the local Lipschitz constant of the dot-product self-attention block showing its dependence on the attention map distributions. The proposed bou…