3 papers
cs.LG2025
EBind: a practical approach to space binding
Jim Broadbent, Felix Cohen, Frederik Hvilshøj +2
We simplify space binding by focusing on two core components, a single encoder per modality and high-quality data; enabling training state-of-the-art models on a single GPU in a fe…
cs.LG2025
Train on Validation (ToV): Fast data selection with applications to fine-tuning
Ayush Jain, Andrea Montanari, Eren Sasoglu
State-of-the-art machine learning often follows a two-stage process: ~pre-training on large, general-purpose datasets; ~fine-tuning on task-specific data. In fine-tuning…
cs.LG2024
Scaling laws for learning with real and surrogate data
Ayush Jain, Andrea Montanari, Eren Sasoglu
Collecting large quantities of high-quality data can be prohibitively expensive or impractical, and a bottleneck in machine learning. One may instead augment a small set of dat…