3 papers
cs.CV2023
In-Style: Bridging Text and Uncurated Videos with Style Transfer for Text-Video Retrieval
Nina Shvetsova, Anna Kukleva, Bernt Schiele +1
Large-scale noisy web image-text datasets have been proven to be efficient for learning robust vision-language models. However, when transferring them to the task of video retrieva…
cs.CV2023
Preserving Modality Structure Improves Multi-Modal Learning
Swetha Sirnam, Mamshad Nayeem Rizve, Nina Shvetsova +2
Self-supervised learning on large-scale multi-modal datasets allows learning semantically meaningful embeddings in a joint multi-modal representation space without relying on human…
cs.CV2022
Augmentation Learning for Semi-Supervised Classification
Tim Frommknecht, Pedro Alves Zipf, Quanfu Fan +2
Recently, a number of new Semi-Supervised Learning methods have emerged. As the accuracy for ImageNet and similar datasets increased over time, the performance on tasks beyond the…