Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
Movie Facts and Fibs (MF): A Benchmark for Long Movie Understanding
Emmanouil Zaranis, António Farinhas, Saul Santos +28
Despite recent progress in vision-language models (VLMs), holistic understanding of long-form video content remains a significant challenge, partly due to limitations in current be…
cs.CV2018
Comparatives, Quantifiers, Proportions: A Multi-Task Model for the Learning of Quantities from Vision
Sandro Pezzelle, Ionut-Teodor Sorodoc, Raffaella Bernardi
The present work investigates whether different quantification mechanisms (set comparison, vague quantification, and proportional estimation) can be jointly learned from visual sce…
cs.CV2017
FOIL it! Find One mismatch between Image and Language caption
Ravi Shekhar, Sandro Pezzelle, Yauhen Klimovich +4
In this paper, we aim to understand whether current language and vision (LaVi) models truly grasp the interaction between the two modalities. To this end, we propose an extension o…