works on

From the 1 of 6 linked papers with an AI index.

activity
20242026
most citedThe Limits and Potentials of Local SGD for Distributed Heterogeneous Learning with Intermittent Communication

1 citations · 1 across the 1 of their papers we have counts for

collaborators

6 papers

cs.LG20261 cited

The Limits and Potentials of Local SGD for Distributed Heterogeneous Learning with Intermittent Communication

Kumar Kshitij Patel, Margalit Glasgow, Ali Zindari +5

The paper analyzes the theoretical limits of Local SGD for distributed learning with heterogeneous data, showing existing heterogeneity assumptions are insufficient for proving its…

cs.LG2026

From Continual Learning to SGD and Back: Better Rates for Continual Linear Models

Itay Evron, Ran Levinstein, Matan Schliserman +4

We study the common continual learning setup where an overparameterized model is sequentially fitted to a set of jointly realizable tasks. We analyze forgetting, defined as the los…

cs.LG2025

Research Program: Theory of Learning in Dynamical Systems

Elad Hazan, Shai Shalev Shwartz, Nathan Srebro

Modern learning systems increasingly interact with data that evolve over time and depend on hidden internal state. We ask a basic question: when is such a dynamical system learnabl…

cs.LG2025

How Uniform Random Weights Induce Non-uniform Bias: Typical Interpolating Neural Networks Generalize with Narrow Teachers

Gon Buzaglo, Itamar Harel, Mor Shpigel Nacson +3

Background. A main theoretical puzzle is why over-parameterized Neural Networks (NNs) generalize well when trained to zero loss (i.e., so they interpolate the data). Usually, the N…

stat.ML2024

The Implicit Bias of Gradient Descent on Separable Data

Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson +2

We examine gradient descent on unregularized logistic regression problems, with homogeneous linear predictors on linearly separable datasets. We show the predictor converges to the…

cs.LG2024

Provable Tempered Overfitting of Minimal Nets and Typical Nets

Itamar Harel, William M. Hoza, Gal Vardi +3

We study the overfitting behavior of fully connected deep Neural Networks (NNs) with binary weights fitted to perfectly classify a noisy training set. We consider interpolation usi…