5 papers · 1 filter
Flatness and Generalization: Learning Multi-Index Models with Homogeneous Neural Networks
Harsh Vardhan, Hossein Taheri, Arya Mazumdar
A common heuristic used to explain the generalization of first-order gradient methods on non-convex neural networks is that "flat interpolators generalize well" (Hochreiter and Sch…
Collaborative Compressors in Distributed Mean Estimation with Limited Communication Budget
Harsh Vardhan, Arya Mazumdar
Distributed high dimensional mean estimation is a common aggregation routine used often in distributed optimization methods. Most of these applications call for a communication-con…
On the Theory of Continual Learning with Gradient Descent for Neural Networks
Hossein Taheri, Avishek Ghosh, Arya Mazumdar
Continual learning, the ability of a model to adapt to an ongoing sequence of tasks without forgetting earlier ones, is a central goal of artificial intelligence. To better underst…
LocalKMeans: Convergence of Lloyd's Algorithm with Distributed Local Iterations
Harsh Vardhan, Heng Zhu, Avishek Ghosh +1
In this paper, we analyze the classical -means alternating-minimization algorithm, also known as Lloyd's algorithm (Lloyd, 1956), for a mixture of Gaussians in a data-distribute…
Learning and Generalization with Mixture Data
Harsh Vardhan, Avishek Ghosh, Arya Mazumdar
In many, if not most, machine learning applications the training data is naturally heterogeneous (e.g. federated learning, adversarial attacks and domain adaptation in neural net t…