2 papers
math.OC2026
Phases of Muon: When Muon Eclipses SignSGD
Elliot Paquette, Noah Marshall, Lucas Benigni +3
Recently, Muon and related spectral optimizers have demonstrated strong empirical performance as scalable stochastic methods, often outperforming Adam. Yet their behaviour remains…
cs.LG2025
Langevin Soft Actor-Critic: Efficient Exploration through Uncertainty-Driven Critic Learning
Haque Ishfaq, Guangyuan Wang, Sami Nur Islam +1
Existing actor-critic algorithms, which are popular for continuous control reinforcement learning (RL) tasks, suffer from poor sample efficiency due to lack of principled explorati…