collaborators

9 papers

math.FA2026

Representation Costs in Data Science: Foundations and the Quasi-Banach Spaces of Deep Neural Networks

Greg Ongie, Rahul Parhi

We develop a general framework for analyzing representation costs induced by parameter-space regularizers in data-fitting methods. For an arbitrary parametric method, we define its…

cs.LG2026

Lifelong In-Context Learning with Transformers Requires Parametric Forms of Attention

Luke McDermott, Robert W. Heath, Rahul Parhi

Lifelong continual learning remains an obstacle on the path to human-like intelligence. Modern transformers show sparks of intelligence with in-context learning. The quadratic natu…

cs.CL2026

LoLA: Low-Rank Linear Attention With Sparse Caching

Luke McDermott, Robert W. Heath, Rahul Parhi

The per-token cost of transformer inference scales with context length, preventing its application to lifelong in-context learning. Linear attention is an efficient alternative tha…

stat.ML2026

Generalization Below the Edge of Stability: The Role of Data Geometry

Tongtong Liang, Alexander Cloninger, Rahul Parhi +1

Understanding generalization in overparameterized neural networks hinges on the interplay between the data geometry, neural architecture, and training dynamics. In this paper, we t…

stat.ML2026

On the Loss Landscape Geometry of Regularized Deep Matrix Factorization: Uniqueness and Sharpness

Anil Kamber, Rahul Parhi

Weight decay is ubiquitous in training deep neural network architectures. Its empirical success is often attributed to capacity control; nonetheless, our theoretical understanding…

stat.ML2026

Sharpness of Minima in Deep Matrix Factorization

Anil Kamber, Rahul Parhi

Understanding the geometry of the loss landscape near a minimum is key to explaining the implicit bias of gradient-based methods in non-convex optimization problems such as deep ne…