1 citations · 1 across the 1 of their papers we have counts for
1 paper · 1 filter
Sho Takishita, Jay Gala, Abdelrahman Mohamed +2
Many vision-language models (VLMs) that prove very effective at a range of multimodal task, build on CLIP-based vision encoders, which are known to have various limitations. We inv…