6 papers
Can Generalist Agents Automate Data Curation?
Feiyang Kang, Hanze Li, Adam Nguyen +5
Curating training data is among the most consequential yet labor-intensive parts of modern AI development: practitioners iteratively propose, implement, evaluate, and revise data p…
Can Small Training Runs Reliably Guide Data Curation? Rethinking Proxy-Model Practice
Jiachen T. Wang, Tong Wu, Kaifeng Lyu +4
Data teams at frontier AI companies routinely train small proxy models to make critical decisions about pretraining data recipes for full-scale training runs. However, the communit…
A Sustainable AI Economy Needs Data Deals That Work for Generators
Ruoxi Jia, Luis Oala, Wenjie Xiong +4
We argue that the machine learning value chain is structurally unsustainable due to an economic data processing inequality: each state in the data cycle from inputs to model weight…
AutoScale: Scale-Aware Data Mixing for Pre-Training LLMs
Feiyang Kang, Yifan Sun, Bingbing Wen +4
Domain reweighting is an emerging research area aimed at adjusting the relative weights of different data sources to improve the effectiveness and efficiency of LLM pre-training. W…
Boosting Alignment for Post-Unlearning Text-to-Image Generative Models
Myeongseob Ko, Henry Li, Zhun Wang +6
Large-scale generative models have shown impressive image-generation capabilities, propelled by massive data. However, this often inadvertently leads to the generation of harmful o…
Capturing the Temporal Dependence of Training Data Influence
Jiachen T. Wang, Dawn Song, James Zou +2
Traditional data influence estimation methods, like influence function, assume that learning algorithms are permutation-invariant with respect to training data. However, modern tra…