14 papers
Natively Unlearnable Large Language Models
Gaurav R. Ghosal, Pratyush Maini, Aditi Raghunathan
Unlearning aims to remove the influence of specific training data sources, but this has proved challenging because the contributions of different sources are entangled within the m…
20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone
DatologyAI, :, Siddharth Joshi +32
Data curation has shifted the quality-compute frontier for language-model and contrastive image-text pretraining, but its role for vision-language models (VLMs) is far less establi…
The Finetuner's Fallacy: When to Pretrain with Your Finetuning Data
Christina Baek, Ricardo Pio Monti, David Schwab +31
Real-world model deployments demand strong performance on narrow domains where data is often scarce. Typically, practitioners finetune models to specialize them, but this risks ove…
ÃberWeb: Insights from Multilingual Curation for a 20-Trillion-Token Dataset
DatologyAI, :, Aldo Gael Carranza +32
Multilinguality is a core capability for modern foundation models, yet training high-quality multilingual models remains challenging due to uneven data availability across language…
When Should We Introduce Safety Interventions During Pretraining?
Dylan Sam, Sachin Goyal, Pratyush Maini +2
Prior work has shown that safety interventions applied during pretraining, such as removing and rephrasing harmful content, can substantially improve the robustness of the resultin…
DatBench: Discriminative, Faithful, and Efficient VLM Evaluations
DatologyAI, :, Siddharth Joshi +30
Empirical evaluation serves as the primary compass guiding research progress in foundation models. Despite a large body of work focused on training frontier vision-language models…