9 papers
Representation Costs in Data Science: Foundations and the Quasi-Banach Spaces of Deep Neural Networks
Greg Ongie, Rahul Parhi
We develop a general framework for analyzing representation costs induced by parameter-space regularizers in data-fitting methods. For an arbitrary parametric method, we define its…
Lifelong In-Context Learning with Transformers Requires Parametric Forms of Attention
Luke McDermott, Robert W. Heath, Rahul Parhi
Lifelong continual learning remains an obstacle on the path to human-like intelligence. Modern transformers show sparks of intelligence with in-context learning. The quadratic natu…
LoLA: Low-Rank Linear Attention With Sparse Caching
Luke McDermott, Robert W. Heath, Rahul Parhi
The per-token cost of transformer inference scales with context length, preventing its application to lifelong in-context learning. Linear attention is an efficient alternative tha…
Generalization Below the Edge of Stability: The Role of Data Geometry
Tongtong Liang, Alexander Cloninger, Rahul Parhi +1
Understanding generalization in overparameterized neural networks hinges on the interplay between the data geometry, neural architecture, and training dynamics. In this paper, we t…
On the Loss Landscape Geometry of Regularized Deep Matrix Factorization: Uniqueness and Sharpness
Anil Kamber, Rahul Parhi
Weight decay is ubiquitous in training deep neural network architectures. Its empirical success is often attributed to capacity control; nonetheless, our theoretical understanding…
Sharpness of Minima in Deep Matrix Factorization
Anil Kamber, Rahul Parhi
Understanding the geometry of the loss landscape near a minimum is key to explaining the implicit bias of gradient-based methods in non-convex optimization problems such as deep ne…