4 papers
Hierarchical Muon: Tiled Newton-Schulz Updates for Efficient Muon Optimization
Ziyuan Tang, Tianshi Xu, Yousef Saad +1
Muon-type optimizers construct update directions for dense neural-network weights by applying a finite Newton-Schulz map to momentum-gradient matrices. For an matrix,…
Design Criteria for SGD Preconditioners: Local Conditioning, Noise Floors, and Basin Stability
Mitchell Scott, Tianshi Xu, Ziyuan Tang +4
Stochastic Gradient Descent (SGD) often slows in the late stage of training due to anisotropic curvature and gradient noise. We analyze preconditioned SGD in the geometry induced b…
Preconditioned Truncated Single-Sample Estimators for Scalable Stochastic Optimization
Tianshi Xu, Difeng Cai, Hua Huang +2
Many large-scale stochastic optimization algorithms involve repeated solutions of linear systems or evaluations of log-determinants. In these regimes, computing exact solutions is…
Mixed Precision Orthogonalization-Free Projection Methods for Eigenvalue and Singular Value Problems
Tianshi Xu, Zechen Zhang, Jie Chen +2
Mixed-precision arithmetic offers significant computational advantages for large-scale matrix computation tasks, yet preserving accuracy and stability in eigenvalue problems and th…