3 papers
cs.LG2026
Spectral Scaling Laws of Muon
Gagik Magakyan, Pablo Parrilo, Asuman Ozdaglar
Orthonormalized update rules have rapidly become a leading choice of optimizer for training large language models, with recent open-source state-of-the-art models adopting Muon. To…
cs.LG2026
Collaborative and Efficient Fine-tuning: Leveraging Task Similarity
Gagik Magakyan, Amirhossein Reisizadeh, Chanwoo Park +2
Adaptability has been regarded as a central feature in the foundation models, enabling them to effectively acclimate to unseen downstream tasks. Parameter-efficient fine-tuning met…
cs.LG2025
Dion: Distributed Orthonormalized Updates
Kwangjun Ahn, Byron Xu, Natalie Abreu +5
Orthonormalized updates accelerate training, improve stability, and enable robust hyperparameter transfer, but existing methods like Muon rely on dense matrix operations that clash…