5 papers · 1 filter
A Sustainable AI Economy Needs Data Deals That Work for Generators
Ruoxi Jia, Luis Oala, Wenjie Xiong +4
We argue that the machine learning value chain is structurally unsustainable due to an economic data processing inequality: each state in the data cycle from inputs to model weight…
Can Small Training Runs Reliably Guide Data Curation? Rethinking Proxy-Model Practice
Jiachen T. Wang, Tong Wu, Kaifeng Lyu +4
Data teams at frontier AI companies routinely train small proxy models to make critical decisions about pretraining data recipes for full-scale training runs. However, the communit…
Capturing the Temporal Dependence of Training Data Influence
Jiachen T. Wang, Dawn Song, James Zou +2
Traditional data influence estimation methods, like influence function, assume that learning algorithms are permutation-invariant with respect to training data. However, modern tra…
Boosting Alignment for Post-Unlearning Text-to-Image Generative Models
Myeongseob Ko, Henry Li, Zhun Wang +6
Large-scale generative models have shown impressive image-generation capabilities, propelled by massive data. However, this often inadvertently leads to the generation of harmful o…
AutoScale: Scale-Aware Data Mixing for Pre-Training LLMs
Feiyang Kang, Yifan Sun, Bingbing Wen +4
Domain reweighting is an emerging research area aimed at adjusting the relative weights of different data sources to improve the effectiveness and efficiency of LLM pre-training. W…