5 papers
Learning to Recall with Transformers Beyond Orthogonal Embeddings
Nuri Mert Vural, Alberto Bietti, Mahdi Soltanolkotabi +1
Modern large language models (LLMs) excel at tasks that require storing and retrieving knowledge, such as factual recall and question answering. Transformers are central to this ca…
Training Dynamics of Softmax Self-Attention: Fast Global Convergence via Preconditioning
Gautam Goel, Mahdi Soltanolkotabi, Peter Bartlett
We study the training dynamics of gradient descent in a softmax self-attention layer trained to perform linear regression and show that a simple first-order optimization algorithm…
CrispEdit: Low-Curvature Projections for Scalable Non-Destructive LLM Editing
Zarif Ikram, Arad Firouzkouhi, Stephen Tu +2
A central challenge in large language model (LLM) editing is capability preservation: methods that successfully change targeted behavior can quietly game the editing proxy and corr…
Stability properties of gradient flow dynamics for the symmetric low-rank matrix factorization problem
Hesameddin Mohammadi, Mohammad Tinati, Stephen Tu +2
The symmetric low-rank matrix factorization serves as a building block in many learning tasks, including matrix recovery and training of neural networks. However, despite a flurry…
Implicit Balancing and Regularization: Generalization and Convergence Guarantees for Overparameterized Asymmetric Matrix Sensing
Mahdi Soltanolkotabi, Dominik Stöger, Changzhi Xie
Recently, there has been significant progress in understanding the convergence and generalization properties of gradient-based methods for training overparameterized learning model…