1 paper
Madhukar Reddy Vongala, Saurabh Srivastava, Jana Košecká
Vision-language pretraining on large datasets of images-text pairs is one of the main building blocks of current Vision-Language Models. While with additional training, these model…