3 papers
cs.LG2026
Navigating LLM Valley: From AdamW to Memory-Efficient and Matrix-Based Optimizers
Aditya Ranganath
Training large language models requires optimization algorithms that are not only statistically effective, but also computationally and memory efficient at extreme scale. Although…
cs.LG2026
A Unified Measure-Theoretic View of Diffusion, Score-Based, and Flow Matching Generative Models
Aditya Ranganath, Mukesh Singhal
We survey continuous-time generative modeling methods based on transporting a simple reference distribution to a data distribution via stochastic or deterministic dynamics. We pres…
math.OC2025
Symmetric Rank-One Quasi-Newton Methods for Deep Learning Using Cubic Regularization
Aditya Ranganath, Mukesh Singhal, Roummel Marcia
Stochastic gradient descent and other first-order variants, such as Adam and AdaGrad, are commonly used in the field of deep learning due to their computational efficiency and low-…