3 papers
cs.LG2026
Scale Dependent Data Duplication
Joshua Kazdan, Noam Levi, Rylan Schaeffer +6
Data duplication during pretraining can degrade generalization and lead to memorization, motivating aggressive deduplication pipelines. However, at web scale, it is unclear what co…
cs.LG2025
Optimal Vector Compressed Sensing Using James Stein Shrinkage
Apratim Dey, David Donoho
The trend in modern science and technology is to take vector measurements rather than scalars, ruthlessly scaling to ever higher dimensional vectors. For about two decades now, tra…
cs.LG2024
Universality of the Pathway in Avoiding Model Collapse
Apratim Dey, David Donoho
Researchers in empirical machine learning recently spotlighted their fears of so-called Model Collapse. They imagined a discard workflow, where an initial generative model is train…