4 papers
TEACHTEXT: CrossModal Generalized Distillation for Text-Video Retrieval
Ioana Croitoru, Simion-Vlad Bogolin, Marius Leordeanu +4
In recent years, considerable progress on the task of text-video retrieval has been achieved by leveraging large-scale pretraining on visual and audio datasets to construct powerfu…
Unsupervised learning of foreground object detection
Ioana Croitoru, Simion-Vlad Bogolin, Marius Leordeanu
Unsupervised learning poses one of the most difficult challenges in computer vision today. The task has an immense practical value with many applications in artificial intelligence…
Mining for meaning: from vision to language through multiple networks consensus
Iulia Duta, Andrei Liviu Nicolicioiu, Simion-Vlad Bogolin +1
Describing visual data into natural language is a very challenging task, at the intersection of computer vision, natural language processing and machine learning. Language goes wel…
Unsupervised learning from video to detect foreground objects in single images
Ioana Croitoru, Simion-Vlad Bogolin, Marius Leordeanu
Unsupervised learning from visual data is one of the most difficult challenges in computer vision, being a fundamental task for understanding how visual recognition works. From a p…