1 paper · 1 filter
Sho Takishita, Jay Gala, Abdelrahman Mohamed +2
Many vision-language models (VLMs) that prove very effective at a range of multimodal task, build on CLIP-based vision encoders, which are known to have various limitations. We inv…