2 papers
cs.CV2026
Dynamic Cluster Data Sampling for Efficient and Long-Tail-Aware Vision-Language Pre-training
Mingliang Liang, Zhuoran Liu, Arjen P. de Vries +1
The computational cost of training a vision-language model (VLM) can be reduced by sampling the training data. Previous work on efficient VLM pre-training has pointed to the import…
cs.IR2026
Tutorial on Reasoning for IR & IR for Reasoning
Mohanna Hoveyda, Panagiotis Efstratiadis, Arjen de Vries +1
Information retrieval has long focused on ranking documents by semantic relatedness. Yet many real-world information needs demand more: enforcement of logical constraints, multi-st…