2 papers
cs.CV2025
Movie Facts and Fibs (MF): A Benchmark for Long Movie Understanding
Emmanouil Zaranis, António Farinhas, Saul Santos +28
Despite recent progress in vision-language models (VLMs), holistic understanding of long-form video content remains a significant challenge, partly due to limitations in current be…
cs.AI2023
GROOViST: A Metric for Grounding Objects in Visual Storytelling
Aditya K Surikuchi, Sandro Pezzelle, Raquel Fernández
A proper evaluation of stories generated for a sequence of images -- the task commonly referred to as visual storytelling -- must consider multiple aspects, such as coherence, gram…