collaborators

14 papers

cs.LG2026

Natively Unlearnable Large Language Models

Gaurav R. Ghosal, Pratyush Maini, Aditi Raghunathan

Unlearning aims to remove the influence of specific training data sources, but this has proved challenging because the contributions of different sources are entangled within the m…

cs.LG2026

20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone

DatologyAI, :, Siddharth Joshi +32

Data curation has shifted the quality-compute frontier for language-model and contrastive image-text pretraining, but its role for vision-language models (VLMs) is far less establi…

cs.LG2026

The Finetuner's Fallacy: When to Pretrain with Your Finetuning Data

Christina Baek, Ricardo Pio Monti, David Schwab +31

Real-world model deployments demand strong performance on narrow domains where data is often scarce. Typically, practitioners finetune models to specialize them, but this risks ove…

cs.LG2026

ÜberWeb: Insights from Multilingual Curation for a 20-Trillion-Token Dataset

DatologyAI, :, Aldo Gael Carranza +32

Multilinguality is a core capability for modern foundation models, yet training high-quality multilingual models remains challenging due to uneven data availability across language…

cs.LG2026

When Should We Introduce Safety Interventions During Pretraining?

Dylan Sam, Sachin Goyal, Pratyush Maini +2

Prior work has shown that safety interventions applied during pretraining, such as removing and rephrasing harmful content, can substantially improve the robustness of the resultin…

cs.LG2026

DatBench: Discriminative, Faithful, and Efficient VLM Evaluations

DatologyAI, :, Siddharth Joshi +30

Empirical evaluation serves as the primary compass guiding research progress in foundation models. Despite a large body of work focused on training frontier vision-language models…