1 citations · 1 across the 2 of their papers we have counts for
1 paper · 1 filter
Maya Varma, Jean-Benoit Delbrouck, Sarah Hooper +2
Vision-language models (VLMs), such as CLIP and ALIGN, are generally trained on datasets consisting of image-caption pairs obtained from the web. However, real-world multimodal dat…