Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Turbo-Muon: Almost-Orthogonal Pre-Conditioning for Fast Muon Updates
Thibaut Boissin, Thomas Massena, Franck Mamalet +1
Orthogonality-based optimizers, such as Muon, have recently shown strong performance across large-scale training and community-driven efficiency challenges. However, these methods…
cs.AI2026
From SGD to Muon: Adaptive Optimization via Schatten-p Norms
Thomas Massena, Corentin Friedrich, Mathieu Serrurier
Modern optimizers, like Muon, impose matrix-wise geometry constraints on their updates. These matrix-wise constraints can be unified under Linear Minimization Oracle (LMO) theory.…
cs.AI2025
An Adaptive Orthogonal Convolution Scheme for Efficient and Flexible CNN Architectures
Thibaut Boissin, Franck Mamalet, Thomas Fel +3
Orthogonal convolutional layers are valuable components in multiple areas of machine learning, such as adversarial robustness, normalizing flows, GANs, and Lipschitz-constrained mo…