collaborators

5 papers

stat.ML2026

Flatness and Generalization: Learning Multi-Index Models with Homogeneous Neural Networks

Harsh Vardhan, Hossein Taheri, Arya Mazumdar

A common heuristic used to explain the generalization of first-order gradient methods on non-convex neural networks is that "flat interpolators generalize well" (Hochreiter and Sch…

stat.ML2026

On the Theory of Continual Learning with Gradient Descent for Neural Networks

Hossein Taheri, Avishek Ghosh, Arya Mazumdar

Continual learning, the ability of a model to adapt to an ongoing sequence of tasks without forgetting earlier ones, is a central goal of artificial intelligence. To better underst…

cs.DC2024

Quantized Decentralized Stochastic Learning over Directed Graphs

Hossein Taheri, Aryan Mokhtari, Hamed Hassani +1

We consider a decentralized stochastic learning problem where data points are distributed among computing nodes communicating over a directed graph. As the model size gets large, d…

cs.LG2024

Sharper Guarantees for Learning Neural Network Classifiers with Gradient Methods

Hossein Taheri, Christos Thrampoulidis, Arya Mazumdar

In this paper, we study the data-dependent convergence and generalization behavior of gradient methods for neural networks with smooth activation. Our first result is a novel bound…

cs.LG2024

On the Optimization and Generalization of Multi-head Attention

Puneesh Deora, Rouzbeh Ghaderi, Hossein Taheri +1

The training and generalization dynamics of the Transformer's core mechanism, namely the Attention mechanism, remain under-explored. Besides, existing analyses primarily focus on s…