activity
20242026
collaborators

7 papers

cs.CV2026

Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding

Arsha Nagrani, Jasper Uijilings, Shyamal Buch +6

Video reasoning models are a core component of egocentric and embodied agents. However, standard benchmarks for assessing models provide only evaluation of the output (e.g. the ans…

cs.CV2026

MINERVA-Cultural: A Benchmark for Cultural and Multilingual Long Video Reasoning

Darshan Singh, Arsha Nagrani, Kawshik Manikantan +6

Recent advancements in video models have shown tremendous progress, particularly in long video understanding. However, current benchmarks predominantly feature western-centric data…

cs.CV2025

CAViAR: Critic-Augmented Video Agentic Reasoning

Sachit Menon, Ahmet Iscen, Arsha Nagrani +3

Video understanding has seen significant progress in recent years, with models' performance on perception from short clips continuing to rise. Yet, multiple recent benchmarks, such…

cs.LG2025

MINERVA: Evaluating Complex Video Reasoning

Arsha Nagrani, Sachit Menon, Ahmet Iscen +9

Multimodal LLMs are turning their focus to video benchmarks, however most video benchmarks only provide outcome supervision, with no intermediate or interpretable reasoning steps.…

cs.LG2025

Neptune: The Long Orbit to Benchmarking Long Video Understanding

Arsha Nagrani, Mingda Zhang, Ramin Mehran +10

We introduce Neptune, a benchmark for long video understanding that requires reasoning over long time horizons and across different modalities. Many existing video datasets and mod…

cs.CV2024

Extending Video Masked Autoencoders to 128 frames

Nitesh Bharadwaj Gundavarapu, Luke Friedman, Raghav Goyal +8

Video understanding has witnessed significant progress with recent video foundation models demonstrating strong performance owing to self-supervised pre-training objectives; Masked…