4 papers
DOSE: Data Selection for Multi-Modal LLMs via Off-the-Shelf Models
Biao Wu, Yiwu Zhong, Meng Fang +1
High-quality and diverse multimodal data are essential for improving vision-language models (VLMs), yet existing datasets often contain noisy, redundant, and poorly aligned samples…
Remote Sensing-Oriented World Model
Yuxi Lu, Biao Wu, Zhidong Li +7
World models have shown potential in artificial intelligence by predicting and reasoning about world states beyond direct observations. However, existing approaches are predominant…
Curriculum Learning with Quality-Driven Data Selection
Biao Wu, Ling Chen
The impressive multimodal capabilities demonstrated by OpenAI's GPT-4 have generated significant interest in the development of Multimodal Large Language Models (MLLMs). Visual ins…
MMCLIP: Cross-modal Attention Masked Modelling for Medical Language-Image Pre-Training
Biao Wu, Yutong Xie, Zeyu Zhang +4
Vision-and-language pretraining (VLP) in the medical field utilizes contrastive learning on image-text pairs to achieve effective transfer across tasks. Yet, current VLP approaches…