activity
20242026
collaborators
Showing 2025Show all

5 papers · 1 filter

cs.CV2025

Filter Like You Test: Data-Driven Data Filtering for CLIP Pretraining

Mikey Shechter, Yair Carmon

We introduce Filter Like You Test (FLYT), an algorithm for curating large-scale vision-language datasets that learns the usefulness of each data point as a pretraining example. FLY…

math.OC2025

Convergence of Clipped SGD on Convex -Smooth Functions

Ofir Gaash, Kfir Yehuda Levy, Yair Carmon

We study stochastic gradient descent (SGD) with gradient clipping on convex functions under a generalized smoothness assumption called -smoothness. Using gradient clippi…

cs.LG2025

DataComp-LM: In search of the next generation of training sets for language models

Jeffrey Li, Alex Fang, Georgios Smyrnis +56

We introduce DataComp for Language Models (DCLM), a testbed for controlled dataset experiments with the goal of improving language models. As part of DCLM, we provide a standardize…

cs.LG2025

An Analytical Model for Overparameterized Learning Under Class Imbalance

Eliav Mor, Yair Carmon

We study class-imbalanced linear classification in a high-dimensional Gaussian mixture model. We develop a tight, closed form approximation for the test error of several practical…

cs.LG2025

Resolving Discrepancies in Compute-Optimal Scaling of Language Models

Tomer Porian, Mitchell Wortsman, Jenia Jitsev +2

Kaplan et al. and Hoffmann et al. developed influential scaling laws for the optimal model size as a function of the compute budget, but these laws yield substantially different pr…