1 paper · 1 filter
Haoran Zhang, Aparna Balagopalan, Nassim Oufattole +4
Large repositories of image-caption pairs are essential for the development of vision-language models. However, these datasets are often extracted from noisy data scraped from the…