5 papers
Flatness and Generalization: Learning Multi-Index Models with Homogeneous Neural Networks
Harsh Vardhan, Hossein Taheri, Arya Mazumdar
A common heuristic used to explain the generalization of first-order gradient methods on non-convex neural networks is that "flat interpolators generalize well" (Hochreiter and Sch…
On the Theory of Continual Learning with Gradient Descent for Neural Networks
Hossein Taheri, Avishek Ghosh, Arya Mazumdar
Continual learning, the ability of a model to adapt to an ongoing sequence of tasks without forgetting earlier ones, is a central goal of artificial intelligence. To better underst…
Quantized Decentralized Stochastic Learning over Directed Graphs
Hossein Taheri, Aryan Mokhtari, Hamed Hassani +1
We consider a decentralized stochastic learning problem where data points are distributed among computing nodes communicating over a directed graph. As the model size gets large, d…
Sharper Guarantees for Learning Neural Network Classifiers with Gradient Methods
Hossein Taheri, Christos Thrampoulidis, Arya Mazumdar
In this paper, we study the data-dependent convergence and generalization behavior of gradient methods for neural networks with smooth activation. Our first result is a novel bound…
On the Optimization and Generalization of Multi-head Attention
Puneesh Deora, Rouzbeh Ghaderi, Hossein Taheri +1
The training and generalization dynamics of the Transformer's core mechanism, namely the Attention mechanism, remain under-explored. Besides, existing analyses primarily focus on s…