2 papers
cs.LG2026
Scaling Muon for Diffusion Transformers
Chenghao Li, Xiao Han, Xinxin Huang +22
The matrix-aware optimizer Muon improves large model training by balancing updates across singular directions, yet its scaling behavior and end-to-end efficiency on large Diffusion…
math.OC2026
AI-Assisted Discovery and Construction of a Counterexample to the Convergence of Three-Block ADMM with the Identity Matrix as its Third Constraint Block
Kenan Xu, Xiangfeng Wang
The alternating direction method of multipliers (ADMM), as a landmark algorithm, has attracted tremendous research attention and extensive practical applications over the past two…