12 papers
Hierarchical Muon: Tiled Newton-Schulz Updates for Efficient Muon Optimization
Ziyuan Tang, Tianshi Xu, Yousef Saad +1
Muon-type optimizers construct update directions for dense neural-network weights by applying a finite Newton-Schulz map to momentum-gradient matrices. For an matrix,…
Factored Sparse Approximate Inverse Preconditioning via Spectral Optimization
Francesco Brarda, Tianshi Xu, Vassilis Kalantzis +2
In this paper, we study value selection for fixed-pattern factorized sparse approximate inverse preconditioners. Given a prescribed sparsity pattern for a factor we choose its…
Hybrid Digital-Analog Approximate Inverse Preconditioning for Krylov Methods
Shikhar Shah, Rui Peng Li, Tayfun Gokmen +3
Analog in-memory computing enables highly parallel matrix-vector multiplications with reduced data movement, but the resulting operations are noisy, quantized, and affected by devi…
A Joint Variational Framework for Multimodal X-ray Ptychography and Fluorescence Reconstruction
Chengru Eric Zou, Elle Buser, Zichao Wendy Di +1
Recovering high-resolution structural and compositional information from coherent X-ray measurements involves solving coupled, nonlinear, and ill-posed inverse problems. Ptychograp…
Design Criteria for SGD Preconditioners: Local Conditioning, Noise Floors, and Basin Stability
Mitchell Scott, Tianshi Xu, Ziyuan Tang +4
Stochastic Gradient Descent (SGD) often slows in the late stage of training due to anisotropic curvature and gradient noise. We analyze preconditioned SGD in the geometry induced b…
Preconditioned Truncated Single-Sample Estimators for Scalable Stochastic Optimization
Tianshi Xu, Difeng Cai, Hua Huang +2
Many large-scale stochastic optimization algorithms involve repeated solutions of linear systems or evaluations of log-determinants. In these regimes, computing exact solutions is…