3 papers
cs.LG2025
An Exploration of Non-Euclidean Gradient Descent: Muon and its Many Variants
Michael Crawshaw, Chirag Modi, Mingrui Liu +1
To define a steepest descent method over a neural network, we need to choose a norm for each layer, a way to aggregate these norms across layers, and whether to use normalization.…
stat.ML2025
Batch, match, and patch: low-rank approximations for score-based variational inference
Chirag Modi, Diana Cai, Lawrence K. Saul
Black-box variational inference (BBVI) scales poorly to high-dimensional problems when it is used to estimate a multivariate Gaussian approximation with a full covariance matrix. I…
stat.ML2024
EigenVI: score-based variational inference with orthogonal function expansions
Diana Cai, Chirag Modi, Charles C. Margossian +3
We develop EigenVI, an eigenvalue-based approach for black-box variational inference (BBVI). EigenVI constructs its variational approximations from orthogonal function expansions.…