4 papers
Hierarchical Muon: Tiled Newton-Schulz Updates for Efficient Muon Optimization
Ziyuan Tang, Tianshi Xu, Yousef Saad +1
Muon-type optimizers construct update directions for dense neural-network weights by applying a finite Newton-Schulz map to momentum-gradient matrices. For an matrix,…
Design Criteria for SGD Preconditioners: Local Conditioning, Noise Floors, and Basin Stability
Mitchell Scott, Tianshi Xu, Ziyuan Tang +4
Stochastic Gradient Descent (SGD) often slows in the late stage of training due to anisotropic curvature and gradient noise. We analyze preconditioned SGD in the geometry induced b…
Toward a Graph Foundation Model: Pre-Training Transformers With Random Walks
Ziyuan Tang, Jie Chen
A foundation model like GPT elicits many emergent abilities, owing to the pre-training with broad inclusion of data and the use of the powerful Transformer architecture. While foun…
Anderson Acceleration with Truncated Gram-Schmidt
Ziyuan Tang, Tianshi Xu, Huan He +2
Anderson Acceleration (AA) is a popular algorithm designed to enhance the convergence of fixed-point iterations. In this paper, we introduce a variant of AA based on a Truncated Gr…