8 papers
Clipping the Price of Adaptivity at the Tail
Itai Kreisler, Yair Carmon, Oliver Hinder
Adaptive stochastic convex optimization (SCO) methods face a fundamental ``price of adaptivity'' barrier: under the standard set of assumptions, they cannot efficiently adapt to la…
The Sample Complexity of Parameter-Free Stochastic Convex Optimization
Jared Lawrence, Ari Kalinsky, Hannah Bradfield +2
We study the sample complexity of stochastic convex optimization when problem parameters such as the distance to optimality and the Lipschitz constant are unknown. We pursue two st…
Filter Like You Test: Data-Driven Data Filtering for CLIP Pretraining
Mikey Shechter, Yair Carmon
We introduce Filter Like You Test (FLYT), an algorithm for curating large-scale vision-language datasets that learns the usefulness of each data point as a pretraining example. FLY…
Convergence of Clipped SGD on Convex -Smooth Functions
Ofir Gaash, Kfir Yehuda Levy, Yair Carmon
We study stochastic gradient descent (SGD) with gradient clipping on convex functions under a generalized smoothness assumption called -smoothness. Using gradient clippi…
DataComp-LM: In search of the next generation of training sets for language models
Jeffrey Li, Alex Fang, Georgios Smyrnis +56
We introduce DataComp for Language Models (DCLM), a testbed for controlled dataset experiments with the goal of improving language models. As part of DCLM, we provide a standardize…
An Analytical Model for Overparameterized Learning Under Class Imbalance
Eliav Mor, Yair Carmon
We study class-imbalanced linear classification in a high-dimensional Gaussian mixture model. We develop a tight, closed form approximation for the test error of several practical…