2 citations · 2 across the 1 of their papers we have counts for
1 paper · 1 filter
Xiaomin Yu, Wenjie Zhang, Ziyue Qiao +2
Training vision-language models (VLMs) typically requires large-scale, high-quality image-text pairs, but collecting or synthesizing such data is costly. In contrast, text data is…