Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Approximate Muon with low-rank adapters
Ben Anson, Conor Houghton, Edward Milsom
The Muon optimizer shows clear benefits versus alternatives when pretraining neural networks. However, it is used less frequently for parameter-efficient fine-tuning (PEFT). One po…
cs.LG2026
Sharpness-Aware Minimization and Muon: Robustness under the Spectral Norm
Wenzhi Zhong, Edward Milsom, Michael Murray
Sharpness-Aware Minimization (SAM) aims to improve generalization by encouraging insensitivity to small, worst-case parameter perturbations. However, the notion of a "small" pertur…