3 papers
cs.LG2026
Demystifying Manifold Constraints in LLM Pre-training
Kang An, Jiaxiang Li, Donald Goldfarb +1
The empirical success of large language model (LLM) pre-training relies heavily on heuristic stabilization techniques, such as explicit normalization layers and weight decay. While…
math.OC2026
Non-Convex Self-Concordant Functions: Practical Algorithms and Complexity Analysis
Donald Goldfarb, Lexiao Lai, Tianyi Lin +1
We extend the standard notion of self-concordance to non-convex optimization and develop a family of second-order algorithms with global convergence guarantees. In particular, two…
cs.LG2025
ASGO: Adaptive Structured Gradient Optimization
Kang An, Yuxing Liu, Rui Pan +4
Training deep neural networks is a structured optimization problem, because the parameters are naturally represented by matrices and tensors rather than by vectors. Under this stru…