1 citations · 1 across the 2 of their papers we have counts for
1 paper · 1 filter
Madhukar Reddy Vongala, Saurabh Srivastava, Jana Košecká
Vision-language pretraining on large datasets of images-text pairs is one of the main building blocks of current Vision-Language Models. While with additional training, these model…