activity
20172025
collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2025

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

Jonathan Roberts, Mohammad Reza Taesiri, Ansh Sharma +31

Large Multimodal Models (LMMs) exhibit shortfalls when interpreting images and, by some measures, have poorer spatial cognition than young children or animals. Despite this, they a…

cs.CV2021

TEACHTEXT: CrossModal Generalized Distillation for Text-Video Retrieval

Ioana Croitoru, Simion-Vlad Bogolin, Marius Leordeanu +4

In recent years, considerable progress on the task of text-video retrieval has been achieved by leveraging large-scale pretraining on visual and audio datasets to construct powerfu…

cs.CV2018

Unsupervised learning of foreground object detection

Ioana Croitoru, Simion-Vlad Bogolin, Marius Leordeanu

Unsupervised learning poses one of the most difficult challenges in computer vision today. The task has an immense practical value with many applications in artificial intelligence…

cs.CV2018

Mining for meaning: from vision to language through multiple networks consensus

Iulia Duta, Andrei Liviu Nicolicioiu, Simion-Vlad Bogolin +1

Describing visual data into natural language is a very challenging task, at the intersection of computer vision, natural language processing and machine learning. Language goes wel…

cs.CV2017

Unsupervised learning from video to detect foreground objects in single images

Ioana Croitoru, Simion-Vlad Bogolin, Marius Leordeanu

Unsupervised learning from visual data is one of the most difficult challenges in computer vision, being a fundamental task for understanding how visual recognition works. From a p…