collaborators

7 papers

cs.CV2025

CAViAR: Critic-Augmented Video Agentic Reasoning

Sachit Menon, Ahmet Iscen, Arsha Nagrani +3

Video understanding has seen significant progress in recent years, with models' performance on perception from short clips continuing to rise. Yet, multiple recent benchmarks, such…

cs.CV2025

VoCap: Video Object Captioning and Segmentation from Any Prompt

Jasper Uijlings, Xingyi Zhou, Xiuye Gu +5

Understanding objects in videos in terms of fine-grained localization masks and detailed semantic properties is a fundamental task in video understanding. In this paper, we propose…

cs.CV2025

OVFact: Measuring and Improving Open-Vocabulary Factuality for Long Caption Models

Monika Wysoczańska, Shyamal Buch, Anurag Arnab +1

Large vision-language models (VLMs) often struggle to generate long and factual captions. However, traditional measures for hallucination and factuality are not well suited for eva…

cs.CV2025

Continual Learning in Vision-Language Models via Aligned Model Merging

Ghada Sokar, Gintare Karolina Dziugaite, Anurag Arnab +3

Continual learning is conventionally tackled through sequential fine-tuning, a process that, while enabling adaptation, inherently favors plasticity over the stability needed to re…

cs.LG2025

MINERVA: Evaluating Complex Video Reasoning

Arsha Nagrani, Sachit Menon, Ahmet Iscen +9

Multimodal LLMs are turning their focus to video benchmarks, however most video benchmarks only provide outcome supervision, with no intermediate or interpretable reasoning steps.…

cs.CV2025

InteractVLM: 3D Interaction Reasoning from 2D Foundational Models

Sai Kumar Dwivedi, Dimitrije Antić, Shashank Tripathi +4

We introduce InteractVLM, a novel method to estimate 3D contact points on human bodies and objects from single in-the-wild images, enabling accurate human-object joint reconstructi…