6 papers
Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation
DatologyAI, :, Matthew L. Leavitt +8
Inference efficiency is typically pursued by shrinking the model: distillation, pruning, quantization, and sparse routing each lower per-token cost while treating token count as fi…
20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone
DatologyAI, :, Siddharth Joshi +32
Data curation has shifted the quality-compute frontier for language-model and contrastive image-text pretraining, but its role for vision-language models (VLMs) is far less establi…
The Finetuner's Fallacy: When to Pretrain with Your Finetuning Data
Christina Baek, Ricardo Pio Monti, David Schwab +31
Real-world model deployments demand strong performance on narrow domains where data is often scarce. Typically, practitioners finetune models to specialize them, but this risks ove…
DatBench: Discriminative, Faithful, and Efficient VLM Evaluations
DatologyAI, :, Siddharth Joshi +30
Empirical evaluation serves as the primary compass guiding research progress in foundation models. Despite a large body of work focused on training frontier vision-language models…
Luxical: High-Speed Lexical-Dense Text Embeddings
DatologyAI, :, Luke Merrick +31
Frontier language model quality increasingly hinges on our ability to organize web-scale text corpora for training. Today's dominant tools trade off speed and flexibility: lexical…
BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining
DatologyAI, :, Pratyush Maini +28
Recent advances in large language model (LLM) pretraining have shown that simply scaling data quantity eventually leads to diminishing returns, hitting a data wall. In response, th…