3 papers
cs.LG2026
20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone
DatologyAI, :, Siddharth Joshi +32
Data curation has shifted the quality-compute frontier for language-model and contrastive image-text pretraining, but its role for vision-language models (VLMs) is far less establi…
cs.LG2025
TensorSocket: Shared Data Loading for Deep Learning Training
Ties Robroek, Neil Kim Nielsen, Pınar Tözün
Training deep learning models is a repetitive and resource-intensive process. Data scientists often train several models before landing on a set of parameters (e.g., hyper-paramete…
cs.LG2025
Modyn: Data-Centric Machine Learning Pipeline Orchestration
Maximilian Böther, Ties Robroek, Viktor Gsteiger +4
In real-world machine learning (ML) pipelines, datasets are continuously growing. Models must incorporate this new training data to improve generalization and adapt to potential di…