2 papers
cs.CV2025
Oasis: One Image is All You Need for Multimodal Instruction Data Synthesis
Letian Zhang, Quan Cui, Bingchen Zhao +1
The success of multi-modal large language models (MLLMs) has been largely attributed to the large-scale training data. However, the training data of many MLLMs is unavailable due t…
cs.CV2024
Vision Learners Meet Web Image-Text Pairs
Bingchen Zhao, Quan Cui, Hao Wu +3
Many self-supervised learning methods are pre-trained on the well-curated ImageNet-1K dataset. In this work, given the excellent scalability of web data, we consider self-supervise…