9 papers
MINERVA-Cultural: A Benchmark for Cultural and Multilingual Long Video Reasoning
Darshan Singh, Arsha Nagrani, Kawshik Manikantan +6
Recent advancements in video models have shown tremendous progress, particularly in long video understanding. However, current benchmarks predominantly feature western-centric data…
CAViAR: Critic-Augmented Video Agentic Reasoning
Sachit Menon, Ahmet Iscen, Arsha Nagrani +3
Video understanding has seen significant progress in recent years, with models' performance on perception from short clips continuing to rise. Yet, multiple recent benchmarks, such…
VoCap: Video Object Captioning and Segmentation from Any Prompt
Jasper Uijlings, Xingyi Zhou, Xiuye Gu +5
Understanding objects in videos in terms of fine-grained localization masks and detailed semantic properties is a fundamental task in video understanding. In this paper, we propose…
OVFact: Measuring and Improving Open-Vocabulary Factuality for Long Caption Models
Monika WysoczaÅska, Shyamal Buch, Anurag Arnab +1
Large vision-language models (VLMs) often struggle to generate long and factual captions. However, traditional measures for hallucination and factuality are not well suited for eva…
Continual Learning in Vision-Language Models via Aligned Model Merging
Ghada Sokar, Gintare Karolina Dziugaite, Anurag Arnab +3
Continual learning is conventionally tackled through sequential fine-tuning, a process that, while enabling adaptation, inherently favors plasticity over the stability needed to re…
MINERVA: Evaluating Complex Video Reasoning
Arsha Nagrani, Sachit Menon, Ahmet Iscen +9
Multimodal LLMs are turning their focus to video benchmarks, however most video benchmarks only provide outcome supervision, with no intermediate or interpretable reasoning steps.…