174 citations · 751 across the 34 of their papers we have counts for
4 papers · 1 filter
Trained Transformers Learn Linear Models In-Context
Ruiqi Zhang, Spencer Frei, Peter L. Bartlett
Attention-based neural networks such as transformers have demonstrated a remarkable ability to exhibit in-context learning (ICL): Given a short prompt sequence of tokens from an un…
Prediction, Learning, Uniform Convergence, and Scale-sensitive Dimensions
Peter L. Bartlett, Philip M. Long
We present a new general-purpose algorithm for learning classes of -valued functions in a generalization of the prediction model, and prove a general upper bound on the expe…
Benign Overfitting in Linear Classifiers and Leaky ReLU Networks from KKT Conditions for Margin Maximization
Spencer Frei, Gal Vardi, Peter L. Bartlett +1
Linear classifiers and leaky ReLU networks trained by gradient flow on the logistic loss have an implicit bias towards solutions which satisfy the Karush--Kuhn--Tucker (KKT) condit…
The Double-Edged Sword of Implicit Bias: Generalization vs. Robustness in ReLU Networks
Spencer Frei, Gal Vardi, Peter L. Bartlett +1
In this work, we study the implications of the implicit bias of gradient flow on generalization and adversarial robustness in ReLU networks. We focus on a setting where the data co…