collaborators

5 papers

cs.LG2025

Drop-Muon: Update Less, Converge Faster

Kaja Gruntkowska, Yassine Maziane, Zheng Qu +1

Conventional wisdom in deep learning optimization dictates updating all layers at every step-a principle followed by all recent state-of-the-art optimizers such as Muon. In this wo…

math.OC2025

Non-Euclidean Broximal Point Method: A Blueprint for Geometry-Aware Optimization

Kaja Gruntkowska, Peter Richtárik

The recently proposed Broximal Point Method (BPM) [Gruntkowska et al., 2025] offers an idealized optimization framework based on iteratively minimizing the objective function over…

cs.LG2025

Error Feedback for Muon and Friends

Kaja Gruntkowska, Alexander Gaponov, Zhirayr Tovmasyan +1

Recent optimizers like Muon, Scion, and Gluon have pushed the frontier of large-scale deep learning by exploiting layer-wise linear minimization oracles (LMOs) over non-Euclidean n…

cs.LG2025

Gluon: Making Muon & Scion Great Again! (Bridging Theory and Practice of LMO-based Optimizers for LLMs)

Artem Riabinin, Egor Shulgin, Kaja Gruntkowska +1

Recent developments in deep learning optimization have brought about radically new algorithms based on the Linear Minimization Oracle (LMO) framework, such as and $\sf S…

math.OC2025

The Ball-Proximal (="Broximal") Point Method: a New Algorithm, Convergence Theory, and Applications

Kaja Gruntkowska, Hanmin Li, Aadi Rane +1

Non-smooth and non-convex global optimization poses significant challenges across various applications, where standard gradient-based methods often struggle. We propose the Ball-Pr…