Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
DOSE: Data Selection for Multi-Modal LLMs via Off-the-Shelf Models
Biao Wu, Yiwu Zhong, Meng Fang +1
High-quality and diverse multimodal data are essential for improving vision-language models (VLMs), yet existing datasets often contain noisy, redundant, and poorly aligned samples…
cs.CV2025
MMCLIP: Cross-modal Attention Masked Modelling for Medical Language-Image Pre-Training
Biao Wu, Yutong Xie, Zeyu Zhang +4
Vision-and-language pretraining (VLP) in the medical field utilizes contrastive learning on image-text pairs to achieve effective transfer across tasks. Yet, current VLP approaches…